In September 2026, the memory price surge made the 8GB band big again. Prices first: per-GB cost from Danawa median actual listings —3060 12GB dominates #1 in the 45,000 KRW range, and the three new 8GB cards run 75,000~100,₩000/GB. If you already own an 8GB card, the table below will get you by.
🌐 · English · 中文 · Español · 한국어 · · English Hub
| Card (median actual sale, Danawa) | Today’s price | per GB |
|---|---|---|
| RTX 3060 12GB | 540,300 KRW | 45,025 won |
| RTX 5050 8GB | ₩605,200 | 75,650 KRW |
| RTX 5060 8GB | 757,000 won | ₩94,625 |
| RTX 5060 Ti 8GB | 801,200 KRW | ₩100,150 |
Builds that actually fit in 8GB
| Model | Build | Weights | Notes |
|---|---|---|---|
| Llama-3.1-8B | Q4_K_M | 4.9 GB | De facto industry default — 6GB line with KV included |
| DeepSeek-R1-Distill-8B | Q4_K_M | 4.9 GB | Inference-specialized |
| Qwen3.5-9B | Q4_K_M | 6.8 GB | The 9B line in the sand — context at 8K |
| Qwen3.5-9B | IQ2_M | 4.9 GB | In a hurry — quality degradation starts |
| Phi-4 | Q3_K_M | 7.2 GB | 14B by force — if it doesn’t fit, IQ2 |
| Llama-3.1-8B | IQ4_XS | 4.4 GB | Saving 0.5GB over Q4 |
3 traps
- Default context: leave the Ollama slider at 256K and a 2GB model balloons to an 18GB footprint and spills onto the CPU — a measured 3.4x slowdown (techfuelhq). Check the CPU/GPU split with ollama ps.
- Automatic GPU offload: LM Studio ‘auto’ quietly leaves part of the MoE on the CPU — raising offload to the max gets 23~31% back (techfuelhq, measured on 5080).
- Don’t shrink KV: just enabling a q8_0 KV cache saves 0.7GB on an 8B basis — a bigger win than using IQ2.
The 12GB-band math isVRAM guide by use, laptopsTGP guide.
Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from sourced public measurements and my own · no local runs ·View the full hub · Quantization ladder, 104 builds · Generation-speed chart, 47 models