‘Which runtime is faster’ is mostly the wrong question — all three run the same llama.cpp engine on the same GGUF file. Differences come only from the packaging (bundled version, defaults, scheduler). Two measured datasets under identical-file sha256 verification are reproduced as-is (inventivehq 2026, techfuelhq 2026-08-26).
🌐 · English · 中文 · Español · 한국어 · · English Hub
| Contender | RTX 5060 Ti (Qwen2.5-Coder-7B Q4) | Apple M3 Max (same file) | Standby RAM |
|---|---|---|---|
| llama.cpp (llama-server) | 77.0 tok/s | 53.5 tok/s | ~75 MB |
| LM Studio | 76.8 (−0.3%, error) | 38.2 (−29%, outdated bundled engine) | 300~500 MB |
| Ollama | 69.1 (−10.3%) | 46.2 (−14%) | 200~400 MB |
One setting outweighs runtime differences
- LM Studio ‘auto’ GPU offload: quietly left gpt-oss 20B on the CPU for up to 31% loss — just maxing the slider takes it from 186→244.9 tok/s (techfuelhq, RTX 5080).
- Ollama context slider: leaving the default 4K at 256K pushes a 3B model onto the CPU, a 3.4× slowdown. Check with ollama ps.
- Quantization choice: Q4_K_M→Q5_K_M eats 15~20% — twice the variable of the 3~8% runtime gap (aibytes aggregate).
Conclusions by situation
- GUI exploration · model experiments: LM Studio — free on NVIDIA (0.3%); AMD Vulkan +12% vs Ollama ROCm.
- API · Docker · background: Ollama — 5~11% ahead on dense decoding, with 5~12x lighter idle RAM. Just check the context default.
- Latest models / maximum control / servers: llama.cpp — fastest upstream support, and the only tool that measures prefill/decode separately with llama-bench.
- Apple Silicon: MLX is home base — in the public M4 Max comparison MLX is +20–25% over GGUF Metal (52.7 vs 43.2). From Ollama v0.40.0 (2026-09-25) MLX is the default, and LM Studio supported it all along. Discard the spring-era benchmark tables.
The Mac buying guide isModels by Mac unified memory, the 5060 Ti practical table is5060 Ti page.
Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from sourced public measurements and my own · no local runs ·View the full hub · Quantization ladder, 104 builds · Generation-speed chart, 47 models