🌐 English · 中文 · Español · 한국어
The gateway that ties every local-LLM page on this site to one hardware question. Pick your spec, get routed below.
| Your situation | Start here |
|---|---|
| 47-model generation-speed chart — main hub — 10 of my own measurements + 6 public benchmarks, newest releases first | 47-model speed chart (KR) |
| Quantization ladder, 104 builds — measured file size per bit-width per model + quality notes | Quantization ladder (KR) |
| Every model that runs on an RTX 5060 Ti 16GB — 13–14.5GB weight budget, 9 measured builds | RTX 5060 Ti 16GB model list (KR) |
| Models by Mac unified memory — budget tiers for 16/32/64/128GB + MLX | Mac unified-memory guide (KR) |
| 8GB GPU survival table — best $/GB is the 3060 12GB + config traps | 8GB GPU survival (KR) |
| Electricity vs Claude API — break-even at ~7–8M output tokens/month | Power cost vs API (KR) |
| Used RTX 3090 cost per token — ₩2.62 per 1,000 tokens + cheapest 48GB path | 3090 token pricing (KR) |
| 4090 LoRA fine-tuning — measured: 8B in 1.6h, 70B in 16h | 4090 LoRA benchmarks (KR) |
| Ollama vs LM Studio vs llama.cpp — front-end measurements + 3 traps | Runtime comparison (KR) |
| VRAM guide by use case — required-VRAM formula per workload | VRAM by use case (KR) |
| Laptop GPU TGP guide — why the same model name performs differently | Laptop TGP guide (KR) |
30-second decision tree
- Windows + NVIDIA ≥16GB → 5060 Ti table (4090/5090: same formula, bigger budget)
- Mac → unified-memory table — start with MLX
- 8GB VRAM → survival table — a 3060 12GB upgrade is the cheapest fix
- Unsure about bit-width → quantization ladder
- No runtime picked → 3-way comparison
- Costing it out → power vs API + 3090 pricing
- Want to train → LoRA benchmarks
Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from attributed public benchmarks and my own runs · nothing executed locally · source pages are in Korean; translations in progress