12GB (RTX 3060/4070/5070) owners: here’s what actually fits and runs well, measured.
| Model | tok/s | VRAM | TTFT | Source |
|---|---|---|---|---|
llama3.1:8b |
100.8 | 4.9GB | 0.1s | measured |
deepseek-r1:8b |
79.7 | 5.2GB | 44.4s | measured |
qwen3.5:9b |
73.0 | 6.6GB | 0.1s | measured |
phi4:14b |
45.8 | 9.1GB | 14.8s | measured |
Comfortable zone: 7–9B at Q4/Q5 (5–7GB) leaves 5GB+ for context. The 12GB card dies at 14B Q4 — 9GB weights + context = offload cliff at 8K tokens. If you’re buying today, 12GB is the hardest tier to justify: +4GB (16GB) roughly doubles the usable model list.
Methodology
Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.
👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기