For training alone, a used RTX 3090 24GB wins. As of September 2026 actual listings, the 3090 runs ₩1.5M~1.8M while a new RTX 4060 Ti 16GB is around ₩1.53M — a gap of only about ₩200K — yet public benches show the 3090 is 1.96x faster at QLoRA training and 2.6x faster at LLM inference. But the warranty, power, and thermal risks all sit on the 3090 side, so it’s not a blind “just get the 3090”. Below I publish the actual prices, benches, and the electricity formula in full.
🌐 · English · 中文 · Español · 한국어 · · English Hub
Used RTX 3090 prices — what they’re selling for now
As checked on September 26, 2026, actual listings on Bunjang show wide variance. MSI SUPRIM X and ASUS ROG STRIX units are up at ₩1.8M, while wanted (buy-it) posts sit around the ₩1.2M–₩1.4M line. Joonggonara’s search average displays ₩2.22M, but that’s a number pulled up by listings in the ₩7M range and above, so don’t take it at face value. In my previous postUsed 3090 vs new 5070As covered in , the practical trading range is ₩1.5M~1.8M, and adding ₩40~50K for a thermal-pad rejob, the real total cost is ₩1.55M~1.85M. On the same day, Danawa lists the MSI 4060 Ti Gaming X Slim 16G at ₩1,522,000 (₩10,000 shipping separate). The effective price gap between the two cards is roughly ₩100~300K.
Training bench: the 3090 was 1.96x faster in QLoRA
The most frequently cited measurement is the same-setup comparison on Reddit r/LocalLLaMA: training Llama 2 with 4-bit QLoRA for 1 epoch (same batch size) on a Threadripper machine with 48GB RAM took 468s on a 3090 and 915s on a 4060 Ti 16GB. 915 divided by 468 is 1.96x. Inference points the same way: in Hardware Corner-style llama.cpp measurements, a Q4-quantized 8B model generates at about 87 tok/s on the 3090 vs about 34 tok/s on the 4060 Ti, a 2.6x gap.
| Item | RTX 3090 24GB | RTX 4060 Ti 16GB |
|---|---|---|
| CUDA cores | 10,496 (per token) | 4,352 (Ada) |
| Memory bandwidth | 936 GB/s (384-bit) | 288 GB/s (128-bit) |
| QLoRA 1 epoch | 468s | 915s (1.96x slower) |
| 8B Q4 inference | about 87 tok/s | about 34 tok/s (2.6x slower) |
| FP8 support | None | Yes (4th-gen tensor) |
| TDP | 350W | 165W |
3.25x the bandwidth decides the match
Why does an older-generation card win? The 4060 Ti’s downfall isn’t compute — it’s the 128-bit memory bus. Check the bandwidth with arithmetic: the 3090 is 19.5Gbps x 384-bit / 8 = 936 GB/s, the 4060 Ti is 18Gbps x 128-bit / 8 = 288 GB/s — exactly 3.25x. Token generation is bandwidth-bound, requiring one read of the weights from memory per token, so even the theoretical ceiling at 8B Q4 (~4.9GB) splits: 3090 = 936 / 4.9 = 191 tok/s, 4060 Ti = 288 / 4.9 = 59 tok/s. Training likewise suffers as batch pressure grows, so the 3090 wins with more VRAM for bigger batches and less gradient accumulation. NVDA’s official AI TOPS alone make the 4060 Ti (353) look higher than the 3090 (285) — a trap that flatly contradicts real-world measurements. To first pin down minimum VRAM per use caseVRAM reference table by use casesee it together.
16GB vs 24GB: this is how differently the load line falls
Q4_K_M quantization is about 0.61GB per billion parameters. 8B is 4.9GB, 14B is 8.5GB, 32B is 19.5GB, and on top of that KV cache plus active buffers add 15–20%. On a 16GB card a 32B doesn’t even need calculating — it won’t load; even a 20B-class model barely fits only if shaved to Q3 or lower. The 24GB 3090 is the only mid-low-price consumer-card route that can run a 32B Q4 while managing context, and about 30 tok/s is reported even in real measurements. If you need high-precision Q8-class residency of the same model, 24GB is the safe line. This 8GB gap decides batch size and sequence length itself in training.
Electricity and hidden costs — where the 4060 Ti wins as a generation should
The power difference is clear. At 8 hours of full load per day, monthly consumption is 0.35kW x 240 hours = 84kWh for the 3090 and 39.6kWh for the 4060 Ti — a gap of 44.4kWh. Converted at the residential low-voltage tier-2 marginal rate (not high-voltage; 214.6 KRW + fuel cost adjustment 5 KRW + climate/environment 9 KRW, about 258 KRW/kWh including VAT and levies), that is roughly 11,500 KRW/month, 138,000 KRW/year. But the electricity for a single training epoch is 350W x 468s / 3600 = 45.5Wh = 11.4 KRW on the 3090, and 165W x 915s = 41.9Wh = 10.5 KRW on the 4060 Ti — even 1000 epochs differ by only 900 KRW. On energy efficiency, the 4060 Ti is the modern-era choice. Over 36-month amortization, the 3090 costs (1.65M + 50K)/36 = 47,222 KRW/month vs 1.53M/36 = 42,556 KRW for the 4060 Ti — 4,666 KRW/month. The real risk of the 3090 isn’t electricity but the missing warranty and GDDR6X heat. Before buying, insist on a 10-minute FurMark run and nvidia-smi memory temperature logs.
Conclusions by budget: student · lab · working professional
- Students (under ₩1.5M):A new 4060 Ti 16GB. 3-year warranty, a 550W PSU is enough, and it covers 8B–14B QLoRA training and experimentation. If you’ve never heard of used-card risk, new is the right answer.
- Lab (2 million won):Used 3090 1.55M–1.85M KRW + a 750W+ PSU. 32B inference and batch-up training finish locally. But price variance is large, so it pays to wait for the 1,600,000 KRW-or-below tier, and if you can’t close a deal in that tierLocal LLM installation guide— also factor cloud-hybrid structures like this into the math.
- Office worker (low noise · cooling essential):Running a 350W old flagship in a living-room-PC-class case is impractical because of noise. Without power headroom and cooling, a 4060 Ti or a quietly-finished 5070 Ti 16GB is the realistic pick.
Frequently asked questions
Q. If I buy a used RTX 3090 now, how long will it stay faster than a 4060 Ti?
A. For inference and QLoRA training, the 3090 is 2~2.6x faster at this very moment, and that gap is a hardware bandwidth limit no software update can flip. That said, on newer runtimes using FP8·FP4 paths the gap narrows somewhat.
Q. What wattage PSU do I need?
A. The 3090 needs a 750W+ system PSU with headroom for peak spikes. The 4060 Ti 16GB is fine with 550W and a single 8-pin. If upgrading an existing PC, a PSU swap adds ₩50~100K.
Q. Is 3090 GDDR6X overheating a real problem?
A. Field reports frequently cite memory temperatures over 100 degrees, so thermal pad replacement is effectively mandatory maintenance. At purchase, check memory temperature after 10 minutes of FurMark; if it’s above 95 degrees, add 40K–50K won of rework cost into your price math.
Q. What do I do if I hit OOM at 16GB during training?
A. First, use gradient accumulation to shrink the micro-batch while keeping the effective batch size. If that still doesn’t fit, drop to QLoRA or shorten the sequence length. On a 3090, most cases end before this step.
Sources: as verified 2026-09-26 — live asking prices on Bunjang/Junggonara, Danawa listed prices (26.09), public benchmarks from Hardware Corner and Reddit r/LocalLLaMA, KEPCO’s Q3 2026 tariff table. Prices and market rates fluctuate — check the latest before buying.