GPU 시세 · VRAM · 온디바이스 AI

Ollama vs llama.cpp vs MLX: Measured Pick for Your Hardware (Oct 2026)

· 읽는 시간 5분 · 데일리 딥러닝

One table, then the one-line answer.

Setup Use Why
Apple Silicon MLX On our M5 Max MLX ran gemma4:26b at 88 tok/s; llama.cpp same model ~30% slower
NVIDIA desktop, single card llama.cpp (via LM Studio or raw) Quant selection, KV cache control, no wrapper tax
Quick service / multi-model server Ollama Convenience tax is ~5% now; API + model management worth it for servers
AMD llama.cpp ROCm Works, narrower tooling; RX 9070 XT is the notable exception (cheap 16GB)

The one-line answer: MLX on Mac, llama.cpp on NVIDIA desktop, Ollama when you’re serving models to other programs.

Methodology

Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.

👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기