can it run
5.4 GB needed of 24 GB usable — headroom for context
24GB · Ampere GA102 · model file 4.9GB · distill 8B
computed roofline: 936 GB/s × 0.46–0.76 efficiency window / 4.9GB (Q4_K_M) — bands, never points.
Community check: ~85 tok/s · 8-9B Q4 · 20 runs · inside our computed band — localmaxxing.com
What users report — 5 cited rows:
→ 92 tok/s · Meta Llama 3.1 8B Q4_K_M llama.cpp · source
→ 95.7 tok/s · Meta Llama 3.1 8B Q4_K_M LocalScore median · source
→ 95 tok/s · Qwen 2.5 7B Q4_K_M llama.cpp b3520, 2K ctx · source
→ 108.0 tok/s · Meta Llama 3.1 8B Q4_0 Ollama 0.3.9 (3-4 run avg, runpod sheet) · source
→ 115.3 tok/s · Qwen3 8B Q4_K llama.cpp llama-bench -fa 1, 4K ctx · source
short-context (≤4k) numbers; measured down-scaling at longer contexts: ×0.75 at 16k, ×0.58 at 32k, ×0.40 at 64k, ×0.24 at 128k
→ Tesla P100 (16GB, used) — cheapest catalog machine that runs DEEPSEEK-R1 distill 8B fully ($135)
→ rtx-5090 — the next step up from NVIDIA RTX 3090
→ every model the NVIDIA RTX 3090 can run — full list
→explore the full catalog84 machines indexed · live prices · what each one can run