GPU 시세 · VRAM · 온디바이스 AI

Apple M5 Max 128GB vs RTX Cards — Measured MLX tok/s Table

· 읽는 시간 5분 · 데일리 딥러닝

Apple silicon gets dismissed in LLM threads constantly, usually without measurements. Here is ours: every model that ran on an M5 Max 128GB, MLX/Ollama, sustained tok/s.

Model tok/s VRAM TTFT Source
qwen3-coder:30b 146.2 18GB 4.3s measured
laguna-xs-2.1:latest 126.0 18GB — public record
nemotron-3.5-lightning:30b-a3b 114.9 25GB — public record
gpt-oss:20b 102.9 13GB 14.2s measured
llama3.1:8b 100.8 4.9GB 0.1s measured
qwen3.6:35b-a3b 95.1 23GB 25.5s measured
ornith-1.5:35b 92.6 37GB — public record
gemma4:26b MLX 88.1 52GB 0.1s measured
deepseek-r1:8b 79.7 5.2GB 44.4s measured
qwen3.8:27b 73.2 17.7GB — public record
qwen3.5:9b 73.0 6.6GB 0.1s measured
qwen3.8-flash-next:125b-mlx 59.0 128.5GB — public record
phi4:14b 45.8 9.1GB 14.8s measured
mistral-small:24b 28.0 14GB 18.2s measured
muse-glimmer:30b-mlx 26.6 18GB — public record
deepseek-r1:70b 10.3 43GB 31.0s measured

Where Apple wins: model capacity (52GB models run fully resident, 0.1s TTFT) and power draw. Where NVIDIA wins: raw tok/s on models that fit VRAM (bandwidth-bound). If your target model exceeds 32GB VRAM, Mac is no longer the slow option — it’s the only option under $4,000.

Methodology

Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.

👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기