Benchmarks

NPU benchmark

The NPU benchmark measures the performance of your system's neural processing unit, a dedicated chip designed to accelerate machine learning and AI workloads. The NPU score appears in your Novabench results only when Novabench detects compatible hardware.

What is an NPU?

A neural processing unit (NPU) is a specialized processor optimized for machine learning inference, the process of running trained models to produce predictions or outputs. Unlike CPUs and GPUs, which handle general-purpose computation, NPUs are designed specifically for the matrix math and tensor operations that machine learning models rely on.

NPUs appear in an increasing number of devices, including recent laptops (marketed as "AI PCs" or "Copilot+ PCs"), and Apple Silicon Macs (as the Neural Engine). Applications often use the NPU for features such as image recognition, natural language processing, background noise removal, or camera effects. The NPU is not necessarily more capable than the system's CPU or GPU at running these models, but it is often more power-efficient.

What the NPU test measures

The NPU benchmark tests two aspects of neural processing performance:

Throughput (TOPS)

TOPS stands for Tera Operations Per Second, the number of trillion operations the NPU can perform each second. Higher TOPS indicates a faster NPU that can process AI workloads more quickly. This metric reflects the raw computational capacity of the neural processor.

Novabench measures throughput using SSD-ResNet50, a realistic object detection model. A pool of concurrent inference requests stresses the NPU's ability to sustain concurrent workloads, rather than only single-stream performance. Novabench sizes the pool to the device during a brief calibration phase.

The TOPS that Novabench records are usually lower than the TOPS that vendors advertise. Vendor values are often theoretical maximums that real models cannot reach.

Inference latency

Inference latency measures how long the NPU takes to process a single input through a neural network model. Lower latency means faster response times for AI features that need real-time results, such as live camera effects, voice recognition, and gesture detection.

Novabench measures latency using Selfie Segmentation. This models the real-world scenario of segmenting every frame of 60 fps video. The result reflects average per-frame inference time. Lower is better.

Score calculation

The NPU score combines throughput and latency into a single number. Novabench must successfully measure both metrics for a score to appear. If either measurement fails (for example, if the NPU is not supported or not available), Novabench omits the NPU score from your results.

Throughput (TOPS) and latency (milliseconds per inference) come on different scales. Before it combines them, Novabench applies a per-workload scaling step calibrated against reference hardware. The two contribute equally. An NPU that is fast at sustained inference but slow per single frame produces a balanced score that reflects both behaviors.

Validation

Before each release, Novabench measures NPU score consistency across supported hardware. These measurements cover release drift, test variance, and vendor alignment.

A note on single-score benchmarks

There is no such thing as a fully objective single-number NPU benchmark. Every score reflects choices about which models to run, which precision to use, and which runtime to call into. Different choices produce different scores, and an NPU that wins one benchmark can lose another depending on what each one emphasizes.

Novabench's choices make scores generally useful for comparing AI accelerators. The test uses real inference models that map to workloads modern applications are starting to offload to the NPU. Throughput and latency contribute equally. Each platform uses its official vendor runtime (Core ML, ONNX QNN, OpenVINO, Vitis AI), so each NPU runs through the path it is designed for. Those choices favor general-purpose on-device AI rather than a single specialized workload, and they don't try to capture every model architecture in the headline number.

If your workload doesn't match those assumptions, the headline score does not tell the full story. Novabench reports the per-workload sub-scores alongside the overall score for exactly this reason. A user who runs sustained AI workloads can look at the throughput result. Someone who needs real-time inference for camera effects, voice features, or live captions can look at the latency result. The headline is a summary; the breakdown shows how the NPU handles your work.

Supported hardware

The NPU benchmark runs automatically when Novabench detects a supported neural processing unit. Novabench supports NPUs through several runtime providers:

Platform

NPU runtime

Examples

macOS

Core ML (Neural Engine)

Apple M1, M2, M3, M4 series

Windows (Qualcomm)

ONNX QNN

Snapdragon X Elite, Snapdragon X Plus

Windows (Intel)

OpenVINO

Intel Core Ultra (Meteor Lake, Arrow Lake, Lunar Lake)

Windows (AMD)

ONNX Vitis AI

AMD Ryzen AI series

Note

NPU availability depends on both hardware and driver support. If your device has an NPU but the test does not run, make sure that the latest NPU drivers are installed. On Windows, NPU drivers are often delivered through Windows Update or the vendor's driver utility.

Downloaded runtimes

Most NPU runtimes ship with Novabench. The AMD Ryzen AI runtime does not. Novabench downloads it on first use, verifies it against a checksum, and caches it alongside the other AI test resources. Novabench only fetches it on AMD hardware that can use it, so the installer stays small for everyone else.

The first NPU benchmark on a Ryzen AI system therefore includes a download. Later runs reuse the cached copy. You can review and clear it from the Manage Resources dialog.

Systems without an NPU

If your system does not have a supported NPU, the benchmark skips the NPU test entirely. Novabench calculates your overall Score from the four core components (CPU, GPU, Memory, Storage), with no penalty for the missing NPU score.

The NPU score is a supplemental metric. Systems with and without NPUs are still fully comparable on the core benchmark components.

When the NPU score matters

NPU performance is relevant when you use (or plan to use) AI-powered features that run locally on your device:

  • Windows AI features: Copilot, Windows Studio Effects (background blur, eye contact, auto framing), live captions, and other AI-powered Windows features use the NPU when available
  • macOS AI features: on-device Siri processing, image analysis, and Core ML-based applications can use the Apple Neural Engine
  • Creative applications: photo and video editing tools with AI-powered features (noise reduction, subject selection, upscaling) can offload work to the NPU

For most general computing tasks (web browsing, office work, gaming), the NPU does not play a significant role. The CPU and GPU remain the primary performance determinants for these workloads.

Cross-platform comparability

Novabench runs the same NPU workloads on every supported platform and vendor. The same inference models run through Core ML on Apple Silicon, ONNX QNN on Qualcomm Snapdragon X, OpenVINO on Intel Core Ultra, and ONNX Vitis AI on AMD Ryzen AI. Each runtime is the path the NPU is designed for, so scores compare cleanly across vendors and operating systems.

How the benchmark runs

The NPU benchmark is designed around three priorities: consistent measurement, fair comparison across systems, and a balanced top-level score backed by full per-test detail. Several mechanics support those priorities:

  • Warmup and calibration: each test starts with a warmup phase that brings the NPU to a steady operating state and calibrates workload size for the system.
  • Process isolation: each test runs in its own worker process, separate from the Novabench app. This lets the benchmark control exactly how the workload runs on the NPU. The app's own UI, logging, and sensor sampling do not interfere. It also stops the previous test from influencing the next one.
  • Real inference models: the throughput test runs SSD-ResNet50, an image object detection model. The latency test runs Selfie Segmentation, which models per-frame inference on 60 fps video. Both are workloads that modern applications are starting to offload to the NPU.
  • Measured TOPS, not theoretical: vendors often publish theoretical peak TOPS that real models cannot reach because of memory bandwidth and other constraints. Novabench reports the operations per second your NPU actually delivers while running real models.
  • Vendor runtime coverage: every result captures the NPU model and the runtime used (Core ML, ONNX QNN, OpenVINO, or Vitis AI). When you compare your score to others, you're comparing the same accelerator running through the same path.

Top-level scores are a summary. Novabench always presents per-test results alongside the overall score, so you can weigh individual workloads against your own use case.

Factors affecting NPU scores

Driver versions

NPU drivers are newer and less mature than CPU or GPU drivers. Performance improvements between driver versions can be significant. Keep NPU drivers up to date through your system manufacturer's update tool or Windows Update.

Thermal conditions

NPUs share the thermal envelope with the CPU (and possibly GPU) in most systems.