When choosing a deep-learning GPU, the first question isn’t performance but load capacity. Even on the same budget, VRAM 12GB vs 24GB changes the actual model size that loads. This article lays out a per-use-case VRAM reference table for deep-learning GPUs with tables and arithmetic, and tags each row with an example GPU.
🌐 · English · 中文 · Español · 한국어 · · English Hub
Why VRAM needs differ by use case
Training bottlenecks in the order VRAM capacity → memory bandwidth → cores; for inference, VRAM is simply the model size you can load. For image generation 8GB is effectively the floor; for LLMs, weights and KV cache must fit simultaneously. That’s why the minimum/recommended figures differ by use case.
| Use case | Minimum VRAM | Recommended VRAM | Caveats | Example GPU |
|---|---|---|---|---|
| Image classification training (transfer learning) | 8 GB | 12 GB | Batch size scales with VRAM | RTX 4060 8GB / 5060 Ti 16GB |
| Stable Diffusion / video generation | 8 GB | 16 GB or more | VRAM balloons as resolution rises | RTX 4060 Ti 16GB |
| LLM inference (7~8B quantized) | 8 GB | 12~16 GB | Bandwidth decides token speed | RTX 5070 12GB |
| LLM inference (13~24B 4bit) | 12 GB | 24 GB | Below 16GB, effectively give up | Used RTX 3090 24GB |
| 3D Gaussian Splatting / NeRF | 12 GB | 24 GB | Large scenes need 24GB | RTX 4090 24GB |
How to calculate the VRAM needed to run an LLM locally
For llama.cpp-family Q4_K_M, weights are about 0.61GB per 1B parameters. On Windows, after CUDA overhead, only 88% of nominal VRAM is actually usable: 12GB becomes ~10.6GB, 24GB ~21.1GB. These two numbers produce the table.
| Model size (Q4_K_M) | Weights (B x 0.61) | KV cache headroom | Total required | Minimum card |
|---|---|---|---|---|
| 8B | 4.9 GB | 3.0 GB | 7.9 GB | 12GB (10.6GB usable, comfortable) |
| 12B | 7.3 GB | 3.3 GB | 10.6 GB | 12GB (10.6GB usable, exact fit) |
| 24B | 14.6 GB | 6.5 GB | 21.1 GB | 24GB (21.1GB usable) |
| 70B | 42.7 GB | 4.0 GB | 46.7 GB | 48GB-class or split across 2 cards |
The fork in the road falls between 12B and 24B. A 24B needs 21.1GB total — its weights alone (14.6GB) don’t fit a 12GB card — and a 70B needs 46.7GB before the 42.2GB usable on a single 48GB card, so a single consumer card was never in play. How far the 5070 12GB goes isRTX 5070 12GB VRAM: real-world limitsI worked it out in more detail at
How many GB for Stable Diffusion
512px generation on SD 1.5 runs on 8GB. But once you move to SDXL 1024px with batch 4+, LoRA training, or video generation, the requirement jumps straight to 16GB or more. In this range an 8GB card fails silently by cutting the generation rather than throwing an error, so for image use I recommend starting at 16GB.
Which GPU fits a beginner training PC
If you’re buying new to get into training, the RTX 4060 Ti 16GB is the most reasonable spot. It holds up for SD default-resolution generation and 7B fine-tuning; add 32GB+ of system RAM too. But its bandwidth is 288GB/s, so limits are clear in large-batch training or fast inference. For beginners whose urgent need is load capacity, not compute speed, capacity comes first.
Realistic picks by lab budget
- Low budget: used RTX 3090 24GB — the only way to get 24GB in the same slot for half the budget
- Middle tier: RTX 5060 Ti 16GB (slower but deeper) or RTX 5070 12GB (faster but shallower) — choose by use-case priority
- High budget: RTX 4090 24GB, 5090 once stock runs out — securing 24GB plus 1TB/s-class bandwidth together
2026 prices — when the table’s numbers update
Live as of the fourth week of September 2026:RTX 5070 ₩1,383,800 · RTX 5060 Ti 8GB ₩801,200 · RTX 3060 12GB ₩540,300 (Danawa actual-listing medians · auto-collected daily at 09:00). The VRAM verdicts in the table above stay price-independent; only prices are updated weekly through this paragraph. The weekly per-GB analysis isPrice tracking weeklycontinues at
Frequently asked questions
Does more VRAM make training proportionally faster?
No. Capacity determines how many you can stack; bandwidth and cores determine speed. A used 3090 matches the 4090 on capacity but trails a generation on speed. Read the two axes separately in the table.
Can I fine-tune with a 12GB card?
7B-class LoRA is the ceiling. Pretraining or fine-tuning 13B+ needs room for optimizer state — 24GB — so if training is the goal, start at 3090/4090 class to avoid wasted quotes.
Can the same reference table be applied to laptops?
Don’t apply that directly. Laptop GPUs at the same model name run at 60~80% of desktop performance depending on TGP, and VRAM is often cut down too. If you assume development and inference on the laptop, training on a server, treat 8GB+ VRAM and 32GB+ RAM as the floor.
Sources & assumptions: llama.cpp Q4_K_M weights at 0.61GB per 1B and the 12% CUDA overhead deduction (88% usable) follow the same assumptions as the 5070 hands-on column; price bands are a May 2026 tally based on the Namuwiki RTX 50 article and Danawa; GPU specs are per manufacturer public data. Street prices fluctuate — check the latest rates before buying.
Buying at today’s prices — market rates for 9 GPUs
Danawa actual listings + used-market snippet aggregation · auto-updated daily at 09:00 KST ·DAILY 09:00
| GPU | Today’s median price | Where it sits in the price range | Trend | vs previous day |
|---|---|---|---|---|
| RTX 5090 32GB | 8,699,990 KRW | ▼ -0.0% | ||
| RTX 5080 16GB | 2,599,000 won | ▲ +3.4% | ||
| RTX 5070 Ti 16GB | ₩1,959,000 | ▼ -1.6% | ||
| RX 9070 XT 16GB | 1,501,600 KRW | ▲ +0.2% | ||
| RTX 5070 12GB | 1,383,800 won | ▲ +0.3% | ||
| RTX 5060 Ti 8GB | 801,200 KRW | ▲ +1.2% | ||
| RTX 5060 8GB | 757,000 won | ▲ +3.3% | ||
| RTX 5050 8GB | ₩605,200 | ▲ +5.8% | ||
| RTX 3060 12GB | 540,300 KRW | ▲ +0.2% |
Last updated 2026-09-28 09:02 KST
.