GPU 시세 · VRAM · 온디바이스 AI

In September 2026, the memory price surge made the 8GB band big again. Prices first: per-GB cost from Danawa median actual listings —3060 12GB dominates #1 in the 45,000 KRW range, and the three new 8GB cards run 75,000~100,₩000/GB. If you already own an 8GB card, the table below will get you by.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Card (median actual sale, Danawa) Today’s price per GB
RTX 3060 12GB 540,300 KRW 45,025 won
RTX 5050 8GB ₩605,200 75,650 KRW
RTX 5060 8GB 757,000 won ₩94,625
RTX 5060 Ti 8GB 801,200 KRW ₩100,150

Builds that actually fit in 8GB

Model Build Weights Notes
Llama-3.1-8B Q4_K_M 4.9 GB De facto industry default — 6GB line with KV included
DeepSeek-R1-Distill-8B Q4_K_M 4.9 GB Inference-specialized
Qwen3.5-9B Q4_K_M 6.8 GB The 9B line in the sand — context at 8K
Qwen3.5-9B IQ2_M 4.9 GB In a hurry — quality degradation starts
Phi-4 Q3_K_M 7.2 GB 14B by force — if it doesn’t fit, IQ2
Llama-3.1-8B IQ4_XS 4.4 GB Saving 0.5GB over Q4

3 traps

  • Default context: leave the Ollama slider at 256K and a 2GB model balloons to an 18GB footprint and spills onto the CPU — a measured 3.4x slowdown (techfuelhq). Check the CPU/GPU split with ollama ps.
  • Automatic GPU offload: LM Studio ‘auto’ quietly leaves part of the MoE on the CPU — raising offload to the max gets 23~31% back (techfuelhq, measured on 5080).
  • Don’t shrink KV: just enabling a q8_0 KV cache saves 0.7GB on an 8B basis — a bigger win than using IQ2.

The 12GB-band math isVRAM guide by use, laptopsTGP guide.

Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from sourced public measurements and my own · no local runs ·View the full hub · Quantization ladder, 104 builds · Generation-speed chart, 47 models