Benchmarks
Benchmark methodology
Novabench benchmarks produce scores that are comparable across Windows, macOS, and Linux, and across different hardware configurations. This page covers the shared design principles that apply to every test. For the specific workloads, models, block sizes, and APIs each test uses, see the per-test pages:
Design principles
Novabench combines synthetic throughput tests with workloads that mirror common application patterns.
Synthetic tests such as raw memory bandwidth, GPU compute, and CPU scalar and SIMD operations measure the peak performance capability of each component under controlled conditions. These tests are sensitive to hardware differences and useful for isolating the performance of a specific subsystem.
Other test patterns mix compute and data access patterns that reflect how software uses hardware in practice. The CPU test includes hash and compression operations alongside arithmetic workloads. The GPU test renders a 3D scene with geometry, textures, lighting, and shaders. The storage test uses both sequential and random access patterns that mirror typical file operations.
Together, these tests produce scores that reflect both raw hardware capability and the performance users experience in practice.
How every test runs
The same set of mechanics applies to every benchmark Novabench runs. Each test page describes how these mechanics show up for that specific workload.
- Fixed-duration workloads: each test runs for a fixed duration rather than a fixed amount of work, and Novabench measures how much work the hardware completes in that time. This approach scales naturally across hardware generations and produces predictable test times.
- Warmup and calibration: every test starts with a warmup that brings the hardware to a steady operating state and calibrates workload size for the system. A fast component and a slow one then both produce a meaningful amount of work to compare.
- Process isolation: each test runs in its own worker process, separate from the Novabench app. This lets the benchmark control exactly how the workload runs on the hardware. The app's own UI, logging, and sensor sampling do not interfere. It also stops the previous test from influencing the next one.
- Thermal gaps between tests: a brief cooldown between tests lets the hardware recover. The gap length was chosen empirically to reduce the thermal impact of prior tests on later ones.
- Hardware detection: every result captures the relevant hardware identity (CPU model and topology, GPU model and driver, memory configuration, drive model and interface, NPU runtime). When you compare your score to others, you're comparing against the same hardware in the same configuration.
- Multi-iteration runs: a single iteration measures system capability with low variance. For tighter precision, you can enable multiple iterations. Each workload then runs several times and Novabench aggregates the results. You trade test time for repeatability.
Scoring approach
Each benchmark produces multiple raw measurements on different scales (GFLOPS, MB/s, milliseconds, TOPS). Two shared steps turn those raw numbers into the scores you see:
- Per-workload scaling: Novabench scales each raw result against reference hardware, so values from different workloads sit on comparable magnitudes. This stops one workload from dominating the final score because of raw-magnitude differences.
- Geometric mean: a geometric mean combines the scaled per-workload results. Each workload contributes equally. A component that excels at one type of operation but lags in another produces a balanced score that reflects overall capability rather than a narrow strength.
Each test page documents the workloads it includes and how the scoring choices apply for that component:
- CPU score calculation
- GPU score calculation
- Memory score calculation
- Storage score calculation
- NPU score calculation
A note on single-score benchmarks
There is no such thing as a fully objective single-number benchmark. Every score reflects choices about which workloads to include, how to weight them, and which platforms and APIs to support. Different choices produce different scores, and a system that wins one benchmark can lose another depending on what each one emphasizes.
Novabench's choices make scores generally useful for comparing systems. Each component uses a balanced mix of workloads, and each workload contributes equally in the geometric mean. The platform choices keep results comparable across operating systems. Top-level scores are a summary. Novabench always presents per-test results alongside them, so you can weigh individual workloads against your own use case. Each test page calls out the specific tradeoffs and which sub-scores to look at for which workloads.
Cross-platform comparability
A Novabench Score of 2,000 is designed to mean the same thing whether it was produced on a Windows desktop, a MacBook, or a Linux workstation.
Each test performs the same computational work across platforms. Novabench runs functionally equivalent workloads on all target operating systems and architectures (x64 and ARM64). Some platform differences make exact equivalence impossible, such as instruction sets that exist on one architecture but not another, or graphics APIs that differ between platforms. In those cases, Novabench makes the tests functionally equivalent rather than identical, so the comparison stays meaningful. Each test page documents the specific platform choices it makes.
Validation
Before each release, Novabench measures score consistency across reference hardware and supported platforms. These measurements cover release drift, test variance, and platform alignment.
Sensor data collection
On Plus, Novabench collects sensor data during the benchmark. It captures sensor readings at regular intervals throughout each test and stores them alongside the results.
What sensors are collected
Sensor | Description |
|---|---|
Temperature | CPU and GPU core temperatures in degrees Celsius |
Power draw | Watts consumed by the CPU and GPU during each test phase |
Clock speed | Processor and GPU clock frequencies in MHz or GHz |
How sensor data helps
Sensor data transforms a benchmark score from a single number into a diagnostic tool:
- Temperature trends: a score that drops across iterations, combined with rising temperatures, confirms thermal throttling. The sensor chart shows exactly when throttling begins.
- Power delivery: unusually low power draw during a GPU test can indicate a power supply problem, or a power-saving mode that limits performance.
- Clock speed stability: a processor that maintains its boost clock throughout the test delivers its full rated performance. Fluctuating clock speeds suggest thermal or power constraints.
Result integrity
Beyond the per-test mechanics above, Novabench includes a few system-level behaviors that protect overall result integrity:
- Failure handling: if a test encounters an error (for example, a GPU driver crash during the 3D test), Novabench flags the result. Flagged results do not enter comparison data.
- Skipped test handling: Novabench cleanly omits tests that do not apply to your system (for example, the NPU test on a system without an NPU). Other scores are unaffected.
- Independent test runs: each component benchmark runs independently, so a failure in one test does not affect the others.
Best practices for reliable results
Follow these guidelines to get the most accurate and repeatable benchmark results:
- Close background applications: other applications can compete with the benchmark for system resources. Close them before you run a benchmark.
- Plug in the power adapter: on laptops, battery power typically limits processor and GPU speed. Benchmark on AC power for full performance.
- Set the power plan: use Balanced or High Performance (Windows), or disable Low Power Mode (macOS). This prevents the system from artificially limiting performance.
- Let the system cool: if you just finished an intensive workload, wait a few minutes for temperatures to return to idle. Elevated starting temperatures reduce the available thermal headroom.
Related pages
- Understanding your scores: what scores mean and how to interpret them
- Comparing results: using histograms, percentile rankings, and Explain to put scores in context
- Stress test: validating system stability under sustained load
