GPU 시세 · VRAM · 온디바이스 AI

Bottom line first: for LLM inference alone, the RTX 4090’s edge is a perceived 1.4–1.5x, not the spec sheet’s 2.7x.At today’s prices on 2026-09-27, a new RTX 5070 is in the 1.37M KRW range and a used RTX 4090 in the 3.34M KRW range — with the price gap at 2.4×, you shouldn’t take the spec sheet’s 2.7× ratio at face value. This article splits the difference between the two cards into two axes, ‘compute’ and ‘load,’ and gives a final verdict on monthly running cost including electricity.

🌐 · English · 中文 · Español · 한국어 · · English Hub

The spec-sheet 2.7x — where it’s true and where it’s a lie

FP32 82.58 vs 30.84 TFLOPS — on pure matrix-multiply density alone, the 4090 is indeed 2.7x. With large image-generation batches and compute-bound training iterations, that gap is felt directly. But LLM inference generates one token at a time and must first read all model parameters from memory. When streaming ~4.5GB where a 7B Q4-quantized model resides, the bottleneck is memory bandwidth, not cores. The bandwidth gap is only 672 vs 1,008 — 1.5x.

Item RTX 5070 (new) RTX 4090 (used)
CUDA cores 6,144 16,384 (2.7x)
FP32 performance 30.84 TFLOPS 82.58 TFLOPS (2.7x)
VRAM 12GB GDDR7 24GB GDDR6X (2x)
Memory bandwidth 672 GB/s 1,008 GB/s (1.5x)
Memory bus 192-bit 384-bit
TDP 250W 450W (+80%)
Efficiency (FP32/W) 0.123 0.183 (5070 is 44% more responsive)
Generation · Interface Blackwell · PCIe 5.0 Ada · PCIe 4.0
Today’s market price ₩1.37M range (new model, 3-year warranty) 3.34M KRW range (used, no warranty)

So just how many times faster is 7B inference, really

1,008 GB/s ÷ ~4.5GB model × 30% effective efficiency

672 GB/s ÷ model ~4.5GB × 30% effective efficiency

The inference bottleneck is bandwidth, not cores

7B Q4 inference (single stream) Basis formula
RTX 4090: ceiling ~224, measured perceived ~67 tok/s
RTX 5070: ceiling about 149, real-world feel about 44 tok/s
→ Gap: 1.5x (not the spec sheet’s 2.7x)

The 5070 with GDDR7 narrows the gap further with its generational edge. On llama.cpp-style runtimes with GPU-Direct-type optimizations, the 7B inference gap between the two cards typically shrinks to a perceived 1.3~1.4x. If you computed ‘buy a 4090 and get 2.7x faster’ from spec-sheet numbers alone, redo the math.

The one condition that still justifies buying a 4090

VRAM 12GB vs 24GB — this isn’t a performance gap, it’s the difference between ‘possible and impossible.’ The detailed math is in the [VRAM reference table](https://daily-llm.com/2026-09-25/vram-guide-by-usecase-2609/), but a 32B-class Q4 model needs about 19.5GB. On a 5070 the attempt doesn’t even stand up; on a 4090 it fits. The [measured column](https://daily-llm.com/2026-09-27/rtx-5070-vram-12gb-deep-learning-limit/) that pins down the RTX 5070’s real-world limits reaches the same conclusion. If your main use is work where the fight is ‘will it run,’ not ‘how many times faster’ — long-context inference, fine-tuning at medium batch, agent stacks that keep several models resident at once — go with 24GB.

How much do monthly running costs diverge once electricity is included

Assuming 720 hours (60% utilization after deducting idle time on a 24h/day basis), the table above gives 65,000 KRW/month for the 5070 and 188,000 KRW/month for a used 4090. On top of the 44% efficiency edge, add the used 4090’s depreciation pace relative to its price and the gap widens to 1.47 million KRW a year. Put differently, a used 4090 at 3.34 million KRW is an option that trades about 51 months of electricity + depreciation against a new 5070 — you’d need serious utilization to recoup that ‘2.7x’ in full.

— the 4090 costs 1.47 million KRW more per year

Monthly running cost (720 hours × 60% load) RTX 5070 RTX 4090 used
Electricity (250 KRW/kWh) 27,000 won 48,600 KRW
Depreciation (5070 over 36 months / used 4090 over 24 months) 38,115 won 138,958 KRW
Total About 65,000 KRW/month about ₩188,000/month
Annual gap

Frequently asked questions

Can’t I just buy two and pool the VRAM?

Tensor Parallel can get you a 24GB setup, but on a 5070 without NVLink the layer-split communication overhead caps total throughput at around 1.5~1.8x. Power hits 500W, and heat and case width climb with it. A single 24GB card usually wins on running costs.

Does electricity really differ that much?

A sustained 200W difference over 432 hours is 21.6kWh per month, ₩5,400 at ₩250/kWh — about ₩22,000 under the 60% load assumption in the table above. Add four-season cooling load and the real gap widens further.

What to check when inspecting a used 4090?

More than mining history, watch for port contact faults, fan noise, and re-pasting history. There’s no warranty, so avoid sellers pushing extended-warranty add-ons and use a channel that allows a one-week test run.

Prices update daily — for today’s lowest price and day-over-day change in real timeGPU price dashboardCheck