GPU 시세 · VRAM · 온디바이스 AI

First Local LLM in 30 Minutes — The Checklist That Works (Oct 2026)

· 읽는 시간 5분 · 데일리 딥러닝

The five decisions that take 30 minutes, in order, with the measured numbers that settle them:

Step Decision rule
1. Check VRAM Windows: taskmgr Performance tab. Mac: About This Mac.
2. Pick runtime Mac→MLX/Ollama, NVIDIA→llama.cpp or Ollama, AMD→Ollama. See our measured comparison.
3. Pick model 8GB: qwen3.5:9b (6.6GB, 73 tok/s). 16GB: qwen3-coder:30b class. 24GB+: gpt-oss:20b and up.
4. Pick quant Q4_K_M default. Q3 only if it’s the only way it fits — reasoning degrades first.
5. Set context Start 8K. Watch TTFT: if it jumps above ~3s you’re offloading — drop context or quant.

Every number above is from our measured table (M5 Max 128GB, October 2026). The TTFT rule is the single most useful thing here — it’s how you know your setup is silently falling off a cliff.

Methodology

Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.

👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기