Uncertainty and calibration
An estimate is a calibrated range, not a benchmark promise. The feed records model, hardware, runtime, quant, context, batch behavior, and measured throughput, then drops outliers, separates batch-one runs, and fits per-stack efficiency into p10, p50, and p90 bands. The middle value is what to plan against. This is calibration in the statistical sense and shares only its name with the quantization step that samples activations to find a clipping range. On one specific machine, a measurement still outranks anything here.
Formula
Worked example: The RTX 4090 8B anchor
Physics says 153 tokens per second. Real runs land between 140 and 180, which is 8.5 percent under to 17.6 percent over. A single number in the middle of that spread would be a claim the data does not support, so the interface shows the spread.