Test any local LLM model across hardware and runtimes, measuring performance and efficiency.

Measure key performance metrics, across prompt sizes.

Measure key performance metrics, across prompt sizes.


Measure how long it takes to read a prompt and generate a response, and how that changes as context size increases.

The same model can run on the CPU, on each GPU through CUDA, Vulkan, Metal, or OpenCL, and on a supported NPU through its vendor runtime. Novabench tests every combination you select in a single run and records the device, runtime, and runtime version with each result.

The presets models cover three size classes: Phi-4 Mini 3.8B at 2.5 GB, Gemma 4 12B at 7 GB, and Gemma 4 31B at 17.7 GB, all Q4 quantized. Larger models need more memory and generate more slowly - Novabench checks for compatibility with your hardware prior to testing.
Novabench's AI benchmark is designed around three priorities: consistent measurement, fair comparison across hardware and runtimes, and raw per-test detail instead of a single score.
Inference throughput and latency for a local language model on each device that can run it. Every completed cell reports prefill throughput, decode throughput, and time to first token. GPU cells, and CPU cells up to 4B parameters, also decode with 4,096 tokens of context. Tokens per watt appears wherever the hardware reports power.
Download Novabench free and find out which models your hardware can run, how fast, and at what power cost.