llm.fit — local LLM fit and speed estimator
Estimate local fit and speed across quantization, KV cache, MoE, offload, and multi-GPU interconnects.
Simulation inputs
Hardware
GPU count
Context
Estimated decode
32tps
p10 26p90 39
weights on GPU 17.3 GBKV 1.5 GBfree 3.7 GBbudget 22.5 GB
- Prefill
- 925 tps
- Time to first token
- 4.4 s
- Memory
- 18.8 GB
- 512-token reply
- 20.2 s