As of today (2026-09-27) at live measured prices, the used market has flipped: median new RTX 5070 at 1,399,000 won,
used RTX 3090 at around 1,800,000 won based on Bunjang FE listings — a price tag where the new card is cheaper than the used one.
The only question left isn’t performance but one thing: do you actually need 24GB?
🌐 · English · 中文 · Español · 한국어 · · English Hub
The price table has already flipped — evidence and traps
| Category | RTX 5070 12GB (new) | RTX 3090 24GB (used) |
|---|---|---|
| Measured asking price | ₩1,399,000 | 1,800,000 KRW |
| Source · date | Danawa measured 2026-09-27 | Beonggae market FE listings / community 1.1–2M |
| Warranty | 3-year official-distribution warranty | None — budget for re-repair separately |
| Amortized monthly over 36 months | 38,861 won | 45,833 won (midpoint of 1.5~1.8 million) |
Community price spread is wide. A May post mentions 1.1M KRW and a Quasar Zone September thread says 2M KRW, but the FE listing actually registered on Beonggae Market today is 1.8M KRW. If you buy a 3090, the real total cost is that actual listing price plus 40–50K KRW for used repair and re-thermal work. The 3-year warranty included with the 5070 belongs on the same scale.
Specs: the core gap is 2x, but it’s the capacity gap that decides
| Item | RTX 3090 (2020) | RTX 5070 (2025) |
|---|---|---|
| VRAM | 24GB GDDR6X | 12GB GDDR7 |
| Bandwidth | 936 GB/s | 672 GB/s |
| FP16 Tensor | 35.58 TFLOPS | 126 TFLOPS |
| FP8 floating point | Not supported | Supported (8-bit cache path) |
| NVLink | Supported (2 cards, 48GB) | Not supported (last supported generation) |
| TDP | 350W | 250W |
| Q8 14B (~14GB weights) | Can stay resident | Offload required |
On FP16 tensor throughput the 5070 is 3.5x ahead, but on memory bandwidth the 3090 actually wins 1.39x (936 vs 672). A Q8-quantized 14B model is ~14GB of weights alone — the 5070’s 12GB has no room for it, CPU offload becomes mandatory and speed halves. On the 3090 there’s even room left for the KV cache. There really is a regime where loading, not compute, decides the outcome.
Inference speed · electricity costs — I show the formulas
| Monthly cost (24-hour inference) | RTX 3090 | RTX 5070 |
|---|---|---|
| Monthly electricity (350/250W) | 37,800 KRW | 27,000 won |
| Monthly token generation (Q4 14B) | 40M tok | 59M tok |
| Productivity per electricity cost | 935 won/M | ₩457/M |
| Price-gap payback | — | Electricity alone saves ₩10,800/month |
Formula: electricity = W x 0.6 (average tensor load) x 720 hours x 250 KRW/kWh. Output = t/s x 86,400 x 30 days x 0.6 (utilization). Inference t/s is based on public benchmarks for Q4 14B single-stream (3090: 26, 5070: 38), measured on llama.cpp-family runtimes. The 5070’s electricity per million tokens is 457 KRW, 51% cheaper than the 3090.
NVLink — the 3090’s one killer advantage
Ampere is the last consumer generation with NVLink. Bridging two 3090s gives 48GB, and 32B-class model inference finishes entirely on-GPU without networking. But bandwidth doesn’t combine and the bottleneck stays, so training needs manual sharding setup, and you must budget add-on boards, case, and PSU headroom together. The total cost of a 2-card build (~3.6M KRW) is half a new 5090 (6.56M KRW) — the only current path to 48GB.
Who should buy what
- Inference ≤14B and light fine-tuning only:The new 5070. Cheaper, faster, under warranty, cheaper to power.
- Q8 14B resident; 24GB is the minimum working unit:Used 3090 + 50,000 KRW budget for re-thermal work.
- 32B inference locally:Two NVLink 3090s. Otherwise the paths are a 5090 or the cloud.
- Awkward middle options:If you need 24GB but DIY building feels like a burden, start with a 4060 Ti 16GB and run cloud in parallel.
Frequently asked questions
Q. Is Q8 14B impossible on an RTX 5070?
A. It runs with offloading. But the CPU-RAM round-trip becomes the bottleneck and single-stream speed drops sharply. With Q4 (about 9GB) it stays within 12GB.
Q. What do I check when buying a used 3090?
A. That’s GDDR6X temperature. Any card whose VRAM has crossed 100 degrees needs a thermal pad replacement, and that should go into the budget from the start. Before buying, ask for 10 minutes of FurMark + an nvidia-smi temperature log.
Q. Do two NVLink-linked 3090s give me 48GB training?
A. Inference is possible. Training requires manual sharding strategy setup (tensor or pipeline parallelism), and due to bandwidth limits total throughput saturates at 1.6~1.9x a single card.