Uncertainty and calibration

An estimate is a calibrated range, not a benchmark promise. The feed records model, hardware, runtime, , , batch behavior, and measured throughput, then drops outliers, separates batch-one runs, and fits per-stack efficiency into p10, p50, and p90 bands. The middle value is what to plan against. This is calibration in the statistical sense and shares only its name with the ization step that samples activations to find a . On one specific machine, a measurement still outranks anything here.

Formula

Worked example: The RTX 4090 8B anchor

Physics says 153 tokens per second. Real runs land between 140 and 180, which is 8.5 percent under to 17.6 percent over. A single number in the middle of that spread would be a claim the data does not support, so the interface shows the spread.

Try a model on the calibrated RTX 4090