Prefill vs decode
Prefill processes the prompt in parallel and is mainly compute-bound. Decode produces the reply one token at a time and is mainly memory-bandwidth-bound. Time to first token, or TTFT, comes from prompt length divided by prefill throughput, while the visible reply pace comes from decode throughput.
Formula
prefill tps = MFU × peak FLOPS ÷ (2 × active params) TTFT = prompt tokens ÷ prefill tps
Worked example: An 8B prompt on an RTX 4090
Using the artifact's RTX 4090 value of 165 FP16 TFLOPS, an 8B active model, and 40% MFU from its 35% to 50% consumer range, prefill is about 4,125 tps. A 4,096-token prompt therefore reaches its first generated token in about one second.
0.40 × 165T ÷ (2 × 8B) = 4,125 tps 4,096 ÷ 4,125 ≈ 0.99 s TTFT