One table, then the one-line answer.
| Setup | Use | Why |
|---|---|---|
| Apple Silicon | MLX | On our M5 Max MLX ran gemma4:26b at 88 tok/s; llama.cpp same model ~30% slower |
| NVIDIA desktop, single card | llama.cpp (via LM Studio or raw) | Quant selection, KV cache control, no wrapper tax |
| Quick service / multi-model server | Ollama | Convenience tax is ~5% now; API + model management worth it for servers |
| AMD | llama.cpp ROCm | Works, narrower tooling; RX 9070 XT is the notable exception (cheap 16GB) |
The one-line answer: MLX on Mac, llama.cpp on NVIDIA desktop, Ollama when you’re serving models to other programs.
Methodology
Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.
👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기