GPU 시세 · VRAM · 온디바이스 AI

When choosing a deep-learning GPU, the first question isn’t performance but load capacity. Even on the same budget, VRAM 12GB vs 24GB changes the actual model size that loads. This article lays out a per-use-case VRAM reference table for deep-learning GPUs with tables and arithmetic, and tags each row with an example GPU.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Why VRAM needs differ by use case

Training bottlenecks in the order VRAM capacity → memory bandwidth → cores; for inference, VRAM is simply the model size you can load. For image generation 8GB is effectively the floor; for LLMs, weights and KV cache must fit simultaneously. That’s why the minimum/recommended figures differ by use case.

Use case Minimum VRAM Recommended VRAM Caveats Example GPU
Image classification training (transfer learning) 8 GB 12 GB Batch size scales with VRAM RTX 4060 8GB / 5060 Ti 16GB
Stable Diffusion / video generation 8 GB 16 GB or more VRAM balloons as resolution rises RTX 4060 Ti 16GB
LLM inference (7~8B quantized) 8 GB 12~16 GB Bandwidth decides token speed RTX 5070 12GB
LLM inference (13~24B 4bit) 12 GB 24 GB Below 16GB, effectively give up Used RTX 3090 24GB
3D Gaussian Splatting / NeRF 12 GB 24 GB Large scenes need 24GB RTX 4090 24GB

How to calculate the VRAM needed to run an LLM locally

For llama.cpp-family Q4_K_M, weights are about 0.61GB per 1B parameters. On Windows, after CUDA overhead, only 88% of nominal VRAM is actually usable: 12GB becomes ~10.6GB, 24GB ~21.1GB. These two numbers produce the table.

Model size (Q4_K_M) Weights (B x 0.61) KV cache headroom Total required Minimum card
8B 4.9 GB 3.0 GB 7.9 GB 12GB (10.6GB usable, comfortable)
12B 7.3 GB 3.3 GB 10.6 GB 12GB (10.6GB usable, exact fit)
24B 14.6 GB 6.5 GB 21.1 GB 24GB (21.1GB usable)
70B 42.7 GB 4.0 GB 46.7 GB 48GB-class or split across 2 cards

The fork in the road falls between 12B and 24B. A 24B needs 21.1GB total — its weights alone (14.6GB) don’t fit a 12GB card — and a 70B needs 46.7GB before the 42.2GB usable on a single 48GB card, so a single consumer card was never in play. How far the 5070 12GB goes isRTX 5070 12GB VRAM: real-world limitsI worked it out in more detail at

How many GB for Stable Diffusion

512px generation on SD 1.5 runs on 8GB. But once you move to SDXL 1024px with batch 4+, LoRA training, or video generation, the requirement jumps straight to 16GB or more. In this range an 8GB card fails silently by cutting the generation rather than throwing an error, so for image use I recommend starting at 16GB.

Which GPU fits a beginner training PC

If you’re buying new to get into training, the RTX 4060 Ti 16GB is the most reasonable spot. It holds up for SD default-resolution generation and 7B fine-tuning; add 32GB+ of system RAM too. But its bandwidth is 288GB/s, so limits are clear in large-batch training or fast inference. For beginners whose urgent need is load capacity, not compute speed, capacity comes first.

Realistic picks by lab budget

  • Low budget: used RTX 3090 24GB — the only way to get 24GB in the same slot for half the budget
  • Middle tier: RTX 5060 Ti 16GB (slower but deeper) or RTX 5070 12GB (faster but shallower) — choose by use-case priority
  • High budget: RTX 4090 24GB, 5090 once stock runs out — securing 24GB plus 1TB/s-class bandwidth together

2026 prices — when the table’s numbers update

Live as of the fourth week of September 2026:RTX 5070 ₩1,383,800 · RTX 5060 Ti 8GB ₩801,200 · RTX 3060 12GB ₩540,300 (Danawa actual-listing medians · auto-collected daily at 09:00). The VRAM verdicts in the table above stay price-independent; only prices are updated weekly through this paragraph. The weekly per-GB analysis isPrice tracking weeklycontinues at

Frequently asked questions

Does more VRAM make training proportionally faster?

No. Capacity determines how many you can stack; bandwidth and cores determine speed. A used 3090 matches the 4090 on capacity but trails a generation on speed. Read the two axes separately in the table.

Can I fine-tune with a 12GB card?

7B-class LoRA is the ceiling. Pretraining or fine-tuning 13B+ needs room for optimizer state — 24GB — so if training is the goal, start at 3090/4090 class to avoid wasted quotes.

Can the same reference table be applied to laptops?

Don’t apply that directly. Laptop GPUs at the same model name run at 60~80% of desktop performance depending on TGP, and VRAM is often cut down too. If you assume development and inference on the laptop, training on a server, treat 8GB+ VRAM and 32GB+ RAM as the floor.

Sources & assumptions: llama.cpp Q4_K_M weights at 0.61GB per 1B and the 12% CUDA overhead deduction (88% usable) follow the same assumptions as the 5070 hands-on column; price bands are a May 2026 tally based on the Namuwiki RTX 50 article and Danawa; GPU specs are per manufacturer public data. Street prices fluctuate — check the latest rates before buying.

Buying at today’s prices — market rates for 9 GPUs

Danawa actual listings + used-market snippet aggregation · auto-updated daily at 09:00 KST ·DAILY 09:00

GPU Today’s median price Where it sits in the price range Trend vs previous day
RTX 5090 32GB 8,699,990 KRW
▼ -0.0%
RTX 5080 16GB 2,599,000 won
▲ +3.4%
RTX 5070 Ti 16GB ₩1,959,000
▼ -1.6%
RX 9070 XT 16GB 1,501,600 KRW
▲ +0.2%
RTX 5070 12GB 1,383,800 won
▲ +0.3%
RTX 5060 Ti 8GB 801,200 KRW
▲ +1.2%
RTX 5060 8GB 757,000 won
▲ +3.3%
RTX 5050 8GB ₩605,200
▲ +5.8%
RTX 3060 12GB 540,300 KRW
▲ +0.2%
New = median of Danawa open-market actual listings · Used = mixed community/used-market asking prices · ±5% vs actual purchase price
Last updated 2026-09-28 09:02 KST

.