GPU 시세 · VRAM · 온디바이스 AI

On Mac it’s unified memory, not VRAM. By convention, what Metal can use is about 75% of the total (raisable via iogpu.wired_limit_mb); subtract 1–4GB for KV and 2GB of system overhead and you get the weight budget. Table sizes are HF library measurements; theoretical speed is bandwidth ÷ active weights.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Unified memory Weight budget Representative entry models (build·size) Speed basis
16GB (M4) ~8GB Llama-3.1-8B Q4_K_M 4.9 · Qwen3.5-9B Q4_K_M 6.8 · DeepSeek-R1-8B Q4 4.9 M4 bandwidth 120GB/s — 8B Q4 theoretical 24 tok/s, reported real-world 15–20
32GB (M4 Pro) ~20GB Qwen3.8-27B UD-Q4_K_M 16.5 · gpt-oss-20b Q8_0 12.1 · Qwen3-Coder-30B Q3_K_M 14.7 M4 Pro 273GB/s — 27B Q4 theoretical 16 tok/s, +20% on MLX
64GB (M4 Max) ~44GB Qwen3.8-27B Q8_0 29.0 · Ornith-1.5-35B Q8_0 37.8 · Qwen3-Coder-30B Q8_0 32.5 · Mistral-Small-24B Q8 25.1 M4 Max 546GB/s — 27B Q8 theoretical 19 tok/s (MLX measurements take priority)
128GB (M5 Max, my measurements) ~92GB Flash-Next 125B UD-Q4_K_XL (MLX 128.5GB line) · Qwen3.8-27B Q8_0 · 35B Q8 · 30B Q8 — all My measurements: Qwen3-Coder-30B 146.2 · gpt-oss-20b 102.9 · 35B-A3B 95.1 · Gemma4-26B 88.1 · Flash-Next 59.0 tok/s

Macs use MLX

  • M4 Max public comparison: Qwen3-Coder-32B MLX 52.7 vs GGUF Metal 43.2 tok/s — MLX +22% (bestllmfor, 2026-04).
  • From Ollama v0.40.0 (2026-09-25), supported models on Apple Silicon switch to MLX execution by default. All the numbers based on an Ollama you set up in the spring are outdated.
  • LM Studio has had first-class MLX support for ages — the 262K Flash-Next only fits in 128GB with MLX 4bit + offload (my measured 59.0 tok/s).

Bit choice isQuantization ladderYou can reuse its per-model table as-is. For laptop Macs, estimate speeds lower than the table above to account for heat and throttling.

Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from sourced public measurements and my own · no local runs ·View the full hub · Quantization ladder, 104 builds · Generation-speed chart, 47 models