llm.fit — local LLM fit and speed estimator

Estimate local fit and speed across , , , , and multi-GPU .

Simulation inputs

Hardware
GPU count
Context

Estimated decode

32tps

p10 26p90 39

Fits
weights on GPU 17.3 GBKV 1.5 GBfree 3.7 GBbudget 22.5 GB
Prefill
925 tps
Time to first token
4.4 s
Memory
18.8 GB
512-token reply
20.2 s

Decode speed vs. context

hover to read
calibratedprojectedoffload begins512256k ctx