Apple silicon gets dismissed in LLM threads constantly, usually without measurements. Here is ours: every model that ran on an M5 Max 128GB, MLX/Ollama, sustained tok/s.
| Model | tok/s | VRAM | TTFT | Source |
|---|---|---|---|---|
qwen3-coder:30b |
146.2 | 18GB | 4.3s | measured |
laguna-xs-2.1:latest |
126.0 | 18GB | — | public record |
nemotron-3.5-lightning:30b-a3b |
114.9 | 25GB | — | public record |
gpt-oss:20b |
102.9 | 13GB | 14.2s | measured |
llama3.1:8b |
100.8 | 4.9GB | 0.1s | measured |
qwen3.6:35b-a3b |
95.1 | 23GB | 25.5s | measured |
ornith-1.5:35b |
92.6 | 37GB | — | public record |
gemma4:26b MLX |
88.1 | 52GB | 0.1s | measured |
deepseek-r1:8b |
79.7 | 5.2GB | 44.4s | measured |
qwen3.8:27b |
73.2 | 17.7GB | — | public record |
qwen3.5:9b |
73.0 | 6.6GB | 0.1s | measured |
qwen3.8-flash-next:125b-mlx |
59.0 | 128.5GB | — | public record |
phi4:14b |
45.8 | 9.1GB | 14.8s | measured |
mistral-small:24b |
28.0 | 14GB | 18.2s | measured |
muse-glimmer:30b-mlx |
26.6 | 18GB | — | public record |
deepseek-r1:70b |
10.3 | 43GB | 31.0s | measured |
Where Apple wins: model capacity (52GB models run fully resident, 0.1s TTFT) and power draw. Where NVIDIA wins: raw tok/s on models that fit VRAM (bandwidth-bound). If your target model exceeds 32GB VRAM, Mac is no longer the slow option — it’s the only option under $4,000.
Methodology
Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.
👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기