Prefill vs decode

Prefill processes the prompt in parallel and is mainly compute-bound. Decode produces the reply one token at a time and is mainly -bound. Time to first token, or TTFT, comes from divided by prefill throughput, while the visible reply pace comes from decode throughput.

Formula

prefill tps = MFU × peak FLOPS ÷ (2 × active params)
TTFT = prompt tokens ÷ prefill tps

Worked example: An 8B prompt on an RTX 4090

Using the artifact's RTX 4090 value of 165 FP16 TFLOPS, an 8B active model, and 40% MFU from its 35% to 50% consumer range, prefill is about 4,125 tps. A 4,096-token prompt therefore reaches its first generated token in about one second.

0.40 × 165T ÷ (2 × 8B) = 4,125 tps
4,096 ÷ 4,125 ≈ 0.99 s TTFT
Try a current 20B model with a 4k prompt