The 16GB card debate keeps repeating without numbers. Here is a measured table: every model we ran, its real VRAM footprint, sustained tok/s, and time-to-first-token. Updated October 2026.
| Model | tok/s | VRAM | TTFT | Source |
|---|---|---|---|---|
gpt-oss:20b |
102.9 | 13GB | 14.2s | measured |
llama3.1:8b |
100.8 | 4.9GB | 0.1s | measured |
deepseek-r1:8b |
79.7 | 5.2GB | 44.4s | measured |
qwen3.5:9b |
73.0 | 6.6GB | 0.1s | measured |
phi4:14b |
45.8 | 9.1GB | 14.8s | measured |
mistral-small:24b |
28.0 | 14GB | 18.2s | measured |
Rule of thumb from the data: Q4_K_M 8B-class models hold 75–100 tok/s on a mid 16GB card; anything above 24B parameters needs offload and TTFT jumps 10–40x. If a benchmark chart doesn’t show TTFT, it’s hiding the offload cliff.
Methodology
Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.
👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기