GPU 시세 · VRAM · 온디바이스 AI

As of today (2026-09-27) at live measured prices, the used market has flipped: median new RTX 5070 at 1,399,000 won,
used RTX 3090 at around 1,800,000 won based on Bunjang FE listings — a price tag where the new card is cheaper than the used one.
The only question left isn’t performance but one thing: do you actually need 24GB?

🌐 · English · 中文 · Español · 한국어 · · English Hub

The price table has already flipped — evidence and traps

Category RTX 5070 12GB (new) RTX 3090 24GB (used)
Measured asking price ₩1,399,000 1,800,000 KRW
Source · date Danawa measured 2026-09-27 Beonggae market FE listings / community 1.1–2M
Warranty 3-year official-distribution warranty None — budget for re-repair separately
Amortized monthly over 36 months 38,861 won 45,833 won (midpoint of 1.5~1.8 million)

Community price spread is wide. A May post mentions 1.1M KRW and a Quasar Zone September thread says 2M KRW, but the FE listing actually registered on Beonggae Market today is 1.8M KRW. If you buy a 3090, the real total cost is that actual listing price plus 40–50K KRW for used repair and re-thermal work. The 3-year warranty included with the 5070 belongs on the same scale.

Specs: the core gap is 2x, but it’s the capacity gap that decides

Item RTX 3090 (2020) RTX 5070 (2025)
VRAM 24GB GDDR6X 12GB GDDR7
Bandwidth 936 GB/s 672 GB/s
FP16 Tensor 35.58 TFLOPS 126 TFLOPS
FP8 floating point Not supported Supported (8-bit cache path)
NVLink Supported (2 cards, 48GB) Not supported (last supported generation)
TDP 350W 250W
Q8 14B (~14GB weights) Can stay resident Offload required

On FP16 tensor throughput the 5070 is 3.5x ahead, but on memory bandwidth the 3090 actually wins 1.39x (936 vs 672). A Q8-quantized 14B model is ~14GB of weights alone — the 5070’s 12GB has no room for it, CPU offload becomes mandatory and speed halves. On the 3090 there’s even room left for the KV cache. There really is a regime where loading, not compute, decides the outcome.

Inference speed · electricity costs — I show the formulas

Monthly cost (24-hour inference) RTX 3090 RTX 5070
Monthly electricity (350/250W) 37,800 KRW 27,000 won
Monthly token generation (Q4 14B) 40M tok 59M tok
Productivity per electricity cost 935 won/M ₩457/M
Price-gap payback — Electricity alone saves ₩10,800/month

Formula: electricity = W x 0.6 (average tensor load) x 720 hours x 250 KRW/kWh. Output = t/s x 86,400 x 30 days x 0.6 (utilization). Inference t/s is based on public benchmarks for Q4 14B single-stream (3090: 26, 5070: 38), measured on llama.cpp-family runtimes. The 5070’s electricity per million tokens is 457 KRW, 51% cheaper than the 3090.

NVLink — the 3090’s one killer advantage

Ampere is the last consumer generation with NVLink. Bridging two 3090s gives 48GB, and 32B-class model inference finishes entirely on-GPU without networking. But bandwidth doesn’t combine and the bottleneck stays, so training needs manual sharding setup, and you must budget add-on boards, case, and PSU headroom together. The total cost of a 2-card build (~3.6M KRW) is half a new 5090 (6.56M KRW) — the only current path to 48GB.

Who should buy what

  • Inference ≤14B and light fine-tuning only:The new 5070. Cheaper, faster, under warranty, cheaper to power.
  • Q8 14B resident; 24GB is the minimum working unit:Used 3090 + 50,000 KRW budget for re-thermal work.
  • 32B inference locally:Two NVLink 3090s. Otherwise the paths are a 5090 or the cloud.
  • Awkward middle options:If you need 24GB but DIY building feels like a burden, start with a 4060 Ti 16GB and run cloud in parallel.

Frequently asked questions

Q. Is Q8 14B impossible on an RTX 5070?
A. It runs with offloading. But the CPU-RAM round-trip becomes the bottleneck and single-stream speed drops sharply. With Q4 (about 9GB) it stays within 12GB.

Q. What do I check when buying a used 3090?
A. That’s GDDR6X temperature. Any card whose VRAM has crossed 100 degrees needs a thermal pad replacement, and that should go into the budget from the start. Before buying, ask for 10 minutes of FurMark + an nvidia-smi temperature log.

Q. Do two NVLink-linked 3090s give me 48GB training?
A. Inference is possible. Training requires manual sharding strategy setup (tensor or pipeline parallelism), and due to bandwidth limits total throughput saturates at 1.6~1.9x a single card.