local intelligence

can it run

Can the NVIDIA RTX 3090 run llama3.1 8B?

FULL

5.4 GB needed of 24 GB usable — headroom for context

24GB · Ampere GA102 · model file 4.9GB · 8B

~88–150 tok/s

computed roofline: 936 GB/s × 0.46–0.76 efficiency window / 4.9GB (Q4_K_M) — bands, never points.

Community check: ~85 tok/s · 8-9B Q4 · 20 runs · inside our computed band — localmaxxing.com

What users report — 5 cited rows:

→ 92 tok/s · Meta Llama 3.1 8B Q4_K_M llama.cpp · source

→ 95.7 tok/s · Meta Llama 3.1 8B Q4_K_M LocalScore median · source

→ 95 tok/s · Qwen 2.5 7B Q4_K_M llama.cpp b3520, 2K ctx · source

→ 108.0 tok/s · Meta Llama 3.1 8B Q4_0 Ollama 0.3.9 (3-4 run avg, runpod sheet) · source

→ 115.3 tok/s · Qwen3 8B Q4_K llama.cpp llama-bench -fa 1, 4K ctx · source

short-context (≤4k) numbers; measured down-scaling at longer contexts: ×0.75 at 16k, ×0.58 at 32k, ×0.40 at 64k, ×0.24 at 128k

Run llama3.1 8B on instead

Tesla P100 (16GB, used) — cheapest catalog machine that runs llama3.1 8B fully ($135)

rtx-5090 — the next step up from NVIDIA RTX 3090

every model the NVIDIA RTX 3090 can run — full list

explore the full catalog84 machines indexed · live prices · what each one can run
2026-09-18 · ← all models on the NVIDIA RTX 3090 · fit = weights + 4k KV vs usable memory