GPU 시세 · VRAM · 온디바이스 AI

🌐 English · 中文 · Español · 한국어

The gateway that ties every local-LLM page on this site to one hardware question. Pick your spec, get routed below.

Your situation Start here
47-model generation-speed chart — main hub — 10 of my own measurements + 6 public benchmarks, newest releases first 47-model speed chart (KR)
Quantization ladder, 104 builds — measured file size per bit-width per model + quality notes Quantization ladder (KR)
Every model that runs on an RTX 5060 Ti 16GB — 13–14.5GB weight budget, 9 measured builds RTX 5060 Ti 16GB model list (KR)
Models by Mac unified memory — budget tiers for 16/32/64/128GB + MLX Mac unified-memory guide (KR)
8GB GPU survival table — best $/GB is the 3060 12GB + config traps 8GB GPU survival (KR)
Electricity vs Claude API — break-even at ~7–8M output tokens/month Power cost vs API (KR)
Used RTX 3090 cost per token — ₩2.62 per 1,000 tokens + cheapest 48GB path 3090 token pricing (KR)
4090 LoRA fine-tuning — measured: 8B in 1.6h, 70B in 16h 4090 LoRA benchmarks (KR)
Ollama vs LM Studio vs llama.cpp — front-end measurements + 3 traps Runtime comparison (KR)
VRAM guide by use case — required-VRAM formula per workload VRAM by use case (KR)
Laptop GPU TGP guide — why the same model name performs differently Laptop TGP guide (KR)

30-second decision tree

Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from attributed public benchmarks and my own runs · nothing executed locally · source pages are in Korean; translations in progress