Benchmarks

CPU benchmark

The CPU benchmark measures your processor's computational performance across single-core and multi-core workloads. The test produces two scores. The multi-core score reflects parallel performance across all cores. The single-threaded score reflects per-core speed.

What the CPU test measures

The CPU benchmark runs a series of computational workloads designed to exercise different aspects of processor performance. Each workload runs for a fixed duration, and Novabench measures how much work the processor completes in that time.

Workload types

The test includes several categories of computation, selected to cover both arithmetically intensive workloads and memory-throughput-heavy workloads:

  • Integer operations: scalar integer multiply-add operations, representative of logic, data manipulation, and general application workloads
  • Floating-point operations: scalar fused multiply-add (FMA) chains that measure per-core decimal math throughput, relevant to scientific computing, media encoding, and 3D calculations
  • SIMD (vectorized floating-point): vectorized FMA operations that measure peak floating-point throughput in GFLOPS. Novabench tests both standard-width SIMD SSE on x64 and NEON on ARM (128-bit) for cross-platform comparison. It also tests the peak instruction set using AVX2 (256-bit) on supported processors for informational purposes.
  • Hash (BLAKE3): hashing random data using the BLAKE3 cryptographic hash function. This is a memory-throughput-heavy workload. Performance depends on how efficiently the CPU streams data through its cache hierarchy to the execution units, rather than on pure arithmetic speed.
  • Compression (gzip): parallel gzip compression of sample data. This is an arithmetically intensive workload that exercises the CPU's ability to find and encode patterns under tight computational loops. In multi-threaded mode, Novabench distributes compression work across all available cores.

Each workload type runs in both single-threaded and multi-threaded modes. The single-threaded tests use a single core with processor affinity set, to keep measurements consistent. The multi-threaded tests distribute work across all available cores.

How the score is calculated

Novabench combines the results from five workload types (SIMD, floating-point, integer, compression, and hash) into a single CPU score using a geometric mean. This method prevents any one workload type from dominating. A processor that excels at one type of operation but lags in another produces a balanced score that reflects overall capability rather than a narrow strength.

Each workload produces its raw result on a different scale. SIMD GFLOPS, integer ops/sec, compression throughput, and BLAKE3 hash rate are not directly comparable numbers. Before it combines them, Novabench applies a per-workload scaling step calibrated against reference hardware, so the values sit on comparable magnitudes. The geometric mean then combines them. Each workload contributes equally, so no single workload dominates the final score because of raw magnitude differences.

The same formula runs separately over the multi-threaded results to produce the multi-core score, and over the single-threaded results to produce the single-threaded score. You therefore see both parallel throughput and per-core speed.

Validation

Before each release, Novabench measures CPU score consistency across reference hardware and supported platforms. These measurements cover release drift, test variance, and platform alignment.

A note on single-score benchmarks

There is no such thing as a fully objective single-number CPU benchmark. Every score reflects choices about which workloads to include, how to weight them, and which platforms and instruction sets to support. Different choices produce different scores, and a CPU that wins one benchmark can lose another depending on what each one emphasizes.

Novabench's choices make scores generally useful for comparing CPUs. Those choices are a balanced mix of arithmetic and memory-bound workloads, equal contribution from each in the geometric mean, and 128-bit SIMD as the scored width. The 128-bit width lets x64 and ARM64 results compare cleanly. Those choices favor general-purpose computing (most desktop software is arithmetic-bound or cache-resident), and they don't try to capture every architectural strength in the headline number. Novabench reports wide SIMD beyond 128 bits but does not score it. A chip with exceptional memory bandwidth shows that strength clearly in its hash sub-score. The headline score reflects it as one of five factors alongside the arithmetic workloads.

If your workload doesn't match those assumptions, the headline score does not tell the full story. Novabench reports the per-workload sub-scores alongside the overall score for exactly this reason. A user who runs heavy video encoding can look at the SIMD and floating-point results. Someone who runs large-scale data processing can look at the hash sub-score as a memory-throughput indicator. Someone who compiles code can look at compression and integer. The headline is a summary; the breakdown shows how the CPU handles your work.

Single-core vs. multi-core

The distinction between single-core and multi-core performance matters because different software uses processors differently.

Single-core performance

Single-core speed determines how fast your processor handles tasks that cannot be split across multiple cores.

Multi-core performance

Multi-core performance reflects how well your processor handles parallel workloads. Applications such as video editing, 3D rendering, software compilation, and scientific simulation distribute work across multiple cores. A processor with many cores and strong multi-threaded performance completes these tasks faster.

Cross-platform comparability

Novabench runs the same CPU workloads on every supported platform: Windows, macOS, and Linux, across x64 and ARM64. The computational work is the same everywhere, so scores compare cleanly across operating systems. Some platform differences make exact equivalence impossible, such as instruction sets that exist on one architecture but not another. In those cases, Novabench makes the tests functionally equivalent rather than identical, so the comparison stays meaningful.

How the benchmark runs

The CPU benchmark is designed around three priorities: consistent measurement, fair comparison across systems, and a balanced top-level score backed by full per-test detail. Several mechanics support those priorities:

  • Warmup and calibration: each test starts with a warmup that brings the CPU to a steady operating state and calibrates workload size to the system. A fast chip and a slow chip then both run a meaningful amount of work.
  • Process isolation: each test runs in its own worker process, separate from the Novabench app. This lets the benchmark control exactly how the workload runs on the CPU. The app's own UI, logging, and sensor sampling do not interfere. It also stops the previous test from influencing the next one.
  • Thermal gaps between tests: a brief cooldown between tests lets the CPU recover, to reduce the thermal impact of prior tests on later ones.
  • Hardware detection and CPU topology: every result captures CPU model, frequency, core count and topology, and virtual machine status. When you compare your score to others, you're comparing against the same chip in the same configuration.
  • Multi-iteration runs for precision: the default run time balances speed and precision. For tighter variance, you can enable multiple iterations. Each workload then runs several times and Novabench aggregates the results. You trade test time for repeatability.

Top-level scores are a summary. Novabench always presents per-test results alongside the overall score, so you can weigh individual workloads against your own use case.

Sensor data during the test

On Plus, Novabench collects sensor data while the CPU benchmark runs. Sensor readings include:

  • Temperature: CPU core temperature over the duration of the test
  • Power draw: watts consumed by the processor during each workload phase
  • Clock speed: processor frequency, which shows whether the CPU maintained its boost clock or throttled under sustained load

Sensor data appears alongside your CPU results on the results screen. A temperature spike combined with a clock speed drop is a clear sign of thermal throttling, which directly reduces your score. See sensor monitoring for more on long-term sensor tracking.

Factors affecting CPU scores

Several variables influence your CPU benchmark results beyond the processor hardware itself.

Power and thermal conditions

  • Power plan: on laptops, power-saving modes limit processor speed. Before you run a benchmark, plug in the power adapter. Then use Balanced or High Performance mode (Windows), or disable Low Power Mode (macOS).
  • Thermal headroom: processors boost to higher clock speeds when temperatures allow. Inadequate cooling, blocked vents, or high ambient temperatures cause the CPU to throttle. This reduces scores. Desktop systems with aftermarket coolers typically score higher than the same processor in a thermally constrained laptop.
  • Sustained vs. burst performance: some processors reach high boost clocks briefly but throttle quickly under sustained load. The CPU benchmark runs long enough to reveal this behavior.

System configuration

  • Background processes: other applications can compete with the benchmark for CPU resources. Close unnecessary applications for the most accurate results.
  • OS updates: operating system updates can affect scheduler behavior and CPU performance.

Hardware factors

  • Core count and thread count: more cores and threads directly improve multi-core scores. Hyper-threading (Intel) or SMT (AMD) adds virtual threads that improve throughput for parallelizable workloads, though each virtual thread is less powerful than a physical core.
  • CPU topology: modern processors often use hybrid architectures with different core types. Intel uses performance cores (P-cores) and efficiency cores (E-cores). Apple Silicon uses performance and efficiency clusters. Novabench detects your processor's topology and reports it alongside results (for example, "8P(16T)+4E" for 8 hyper-threaded performance cores plus 4 efficiency cores). All core types contribute to multi-threaded scores, but P-cores deliver significantly more throughput per core than E-cores.
  • Clock speed: higher base and boost clocks improve both single-core and multi-core scores. Boost clock behavior varies by workload duration and thermal conditions.
  • Architecture generation: newer processor architectures generally deliver better performance per clock cycle (IPC). A newer processor at the same clock speed as an older one typically scores higher.