The five decisions that take 30 minutes, in order, with the measured numbers that settle them:
| Step | Decision rule |
|---|---|
| 1. Check VRAM | Windows: taskmgr Performance tab. Mac: About This Mac. |
| 2. Pick runtime | Mac→MLX/Ollama, NVIDIA→llama.cpp or Ollama, AMD→Ollama. See our measured comparison. |
| 3. Pick model | 8GB: qwen3.5:9b (6.6GB, 73 tok/s). 16GB: qwen3-coder:30b class. 24GB+: gpt-oss:20b and up. |
| 4. Pick quant | Q4_K_M default. Q3 only if it’s the only way it fits — reasoning degrades first. |
| 5. Set context | Start 8K. Watch TTFT: if it jumps above ~3s you’re offloading — drop context or quant. |
Every number above is from our measured table (M5 Max 128GB, October 2026). The TTFT rule is the single most useful thing here — it’s how you know your setup is silently falling off a cliff.
Methodology
Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.
👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기