GPU 시세 · VRAM · 온디바이스 AI

Every Local LLM That Fits 16GB — Measured tok/s (Oct 2026)

· 읽는 시간 5분 · 데일리 딥러닝

The 16GB card debate keeps repeating without numbers. Here is a measured table: every model we ran, its real VRAM footprint, sustained tok/s, and time-to-first-token. Updated October 2026.

Model tok/s VRAM TTFT Source
gpt-oss:20b 102.9 13GB 14.2s measured
llama3.1:8b 100.8 4.9GB 0.1s measured
deepseek-r1:8b 79.7 5.2GB 44.4s measured
qwen3.5:9b 73.0 6.6GB 0.1s measured
phi4:14b 45.8 9.1GB 14.8s measured
mistral-small:24b 28.0 14GB 18.2s measured

Rule of thumb from the data: Q4_K_M 8B-class models hold 75–100 tok/s on a mid 16GB card; anything above 24B parameters needs offload and TTFT jumps 10–40x. If a benchmark chart doesn’t show TTFT, it’s hiding the offload cliff.

Methodology

Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.

👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기