GPU 시세 · VRAM · 온디바이스 AI

‘Which runtime is faster’ is mostly the wrong question — all three run the same llama.cpp engine on the same GGUF file. Differences come only from the packaging (bundled version, defaults, scheduler). Two measured datasets under identical-file sha256 verification are reproduced as-is (inventivehq 2026, techfuelhq 2026-08-26).

🌐 · English · 中文 · Español · 한국어 · · English Hub

Contender RTX 5060 Ti (Qwen2.5-Coder-7B Q4) Apple M3 Max (same file) Standby RAM
llama.cpp (llama-server) 77.0 tok/s 53.5 tok/s ~75 MB
LM Studio 76.8 (−0.3%, error) 38.2 (−29%, outdated bundled engine) 300~500 MB
Ollama 69.1 (−10.3%) 46.2 (−14%) 200~400 MB

One setting outweighs runtime differences

  • LM Studio ‘auto’ GPU offload: quietly left gpt-oss 20B on the CPU for up to 31% loss — just maxing the slider takes it from 186→244.9 tok/s (techfuelhq, RTX 5080).
  • Ollama context slider: leaving the default 4K at 256K pushes a 3B model onto the CPU, a 3.4× slowdown. Check with ollama ps.
  • Quantization choice: Q4_K_M→Q5_K_M eats 15~20% — twice the variable of the 3~8% runtime gap (aibytes aggregate).

Conclusions by situation

  • GUI exploration · model experiments: LM Studio — free on NVIDIA (0.3%); AMD Vulkan +12% vs Ollama ROCm.
  • API · Docker · background: Ollama — 5~11% ahead on dense decoding, with 5~12x lighter idle RAM. Just check the context default.
  • Latest models / maximum control / servers: llama.cpp — fastest upstream support, and the only tool that measures prefill/decode separately with llama-bench.
  • Apple Silicon: MLX is home base — in the public M4 Max comparison MLX is +20–25% over GGUF Metal (52.7 vs 43.2). From Ollama v0.40.0 (2026-09-25) MLX is the default, and LM Studio supported it all along. Discard the spring-era benchmark tables.

The Mac buying guide isModels by Mac unified memory, the 5060 Ti practical table is5060 Ti page.

Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from sourced public measurements and my own · no local runs ·View the full hub · Quantization ladder, 104 builds · Generation-speed chart, 47 models