Bottom line first: for LLM inference alone, the RTX 4090’s edge is a perceived 1.4–1.5x, not the spec sheet’s 2.7x.At today’s prices on 2026-09-27, a new RTX 5070 is in the 1.37M KRW range and a used RTX 4090 in the 3.34M KRW range — with the price gap at 2.4×, you shouldn’t take the spec sheet’s 2.7× ratio at face value. This article splits the difference between the two cards into two axes, ‘compute’ and ‘load,’ and gives a final verdict on monthly running cost including electricity.
🌐 · English · 中文 · Español · 한국어 · · English Hub
The spec-sheet 2.7x — where it’s true and where it’s a lie
FP32 82.58 vs 30.84 TFLOPS — on pure matrix-multiply density alone, the 4090 is indeed 2.7x. With large image-generation batches and compute-bound training iterations, that gap is felt directly. But LLM inference generates one token at a time and must first read all model parameters from memory. When streaming ~4.5GB where a 7B Q4-quantized model resides, the bottleneck is memory bandwidth, not cores. The bandwidth gap is only 672 vs 1,008 — 1.5x.
| Item | RTX 5070 (new) | RTX 4090 (used) |
|---|---|---|
| CUDA cores | 6,144 | 16,384 (2.7x) |
| FP32 performance | 30.84 TFLOPS | 82.58 TFLOPS (2.7x) |
| VRAM | 12GB GDDR7 | 24GB GDDR6X (2x) |
| Memory bandwidth | 672 GB/s | 1,008 GB/s (1.5x) |
| Memory bus | 192-bit | 384-bit |
| TDP | 250W | 450W (+80%) |
| Efficiency (FP32/W) | 0.123 | 0.183 (5070 is 44% more responsive) |
| Generation · Interface | Blackwell · PCIe 5.0 | Ada · PCIe 4.0 |
| Today’s market price | ₩1.37M range (new model, 3-year warranty) | 3.34M KRW range (used, no warranty) |
So just how many times faster is 7B inference, really
| 7B Q4 inference (single stream) | Basis formula |
|---|---|
| RTX 4090: ceiling ~224, measured perceived ~67 tok/s | |
| RTX 5070: ceiling about 149, real-world feel about 44 tok/s | |
| → Gap: 1.5x (not the spec sheet’s 2.7x) |
The 5070 with GDDR7 narrows the gap further with its generational edge. On llama.cpp-style runtimes with GPU-Direct-type optimizations, the 7B inference gap between the two cards typically shrinks to a perceived 1.3~1.4x. If you computed ‘buy a 4090 and get 2.7x faster’ from spec-sheet numbers alone, redo the math.
The one condition that still justifies buying a 4090
VRAM 12GB vs 24GB — this isn’t a performance gap, it’s the difference between ‘possible and impossible.’ The detailed math is in the [VRAM reference table](https://daily-llm.com/2026-09-25/vram-guide-by-usecase-2609/), but a 32B-class Q4 model needs about 19.5GB. On a 5070 the attempt doesn’t even stand up; on a 4090 it fits. The [measured column](https://daily-llm.com/2026-09-27/rtx-5070-vram-12gb-deep-learning-limit/) that pins down the RTX 5070’s real-world limits reaches the same conclusion. If your main use is work where the fight is ‘will it run,’ not ‘how many times faster’ — long-context inference, fine-tuning at medium batch, agent stacks that keep several models resident at once — go with 24GB.
How much do monthly running costs diverge once electricity is included
Assuming 720 hours (60% utilization after deducting idle time on a 24h/day basis), the table above gives 65,000 KRW/month for the 5070 and 188,000 KRW/month for a used 4090. On top of the 44% efficiency edge, add the used 4090’s depreciation pace relative to its price and the gap widens to 1.47 million KRW a year. Put differently, a used 4090 at 3.34 million KRW is an option that trades about 51 months of electricity + depreciation against a new 5070 — you’d need serious utilization to recoup that ‘2.7x’ in full.
| Monthly running cost (720 hours × 60% load) | RTX 5070 | RTX 4090 used |
|---|---|---|
| Electricity (250 KRW/kWh) | 27,000 won | 48,600 KRW |
| Depreciation (5070 over 36 months / used 4090 over 24 months) | 38,115 won | 138,958 KRW |
| Total | About 65,000 KRW/month | about ₩188,000/month |
| Annual gap |
Frequently asked questions
Can’t I just buy two and pool the VRAM?
Tensor Parallel can get you a 24GB setup, but on a 5070 without NVLink the layer-split communication overhead caps total throughput at around 1.5~1.8x. Power hits 500W, and heat and case width climb with it. A single 24GB card usually wins on running costs.
Does electricity really differ that much?
A sustained 200W difference over 432 hours is 21.6kWh per month, ₩5,400 at ₩250/kWh — about ₩22,000 under the 60% load assumption in the table above. Add four-season cooling load and the real gap widens further.
What to check when inspecting a used 4090?
More than mining history, watch for port contact faults, fan noise, and re-pasting history. There’s no warranty, so avoid sellers pushing extended-warranty add-ons and use a channel that allows a one-week test run.
Prices update daily — for today’s lowest price and day-over-day change in real timeGPU price dashboardCheck