GPU 시세 · VRAM · 온디바이스 AI

What you can train on a single 4090 24GB, and at what cost — the recipe table is gigagpu’s measured data, with Qwen mappings added to suit my machine.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Recipes Model Sequence Batch Processing speed 50k samples, 1 epoch
LoRA bf16 Llama-3-8B (equivalent to Qwen3-8B) 2048 8 ~18,000 tok/s ~1.6 hours
LoRA bf16 Mistral-7B 2048 8 ~21,000 tok/s ~1.4 hours
LoRA bf16 Qwen2.5-14B 2048 2 ~6,400 tok/s ~4.4 hours
QLoRA NF4 Llama-3-8B 2048 8 ~14,500 tok/s ~2.0 hours
QLoRA NF4 Llama-3-70B 2048 1 ~1,800 tok/s ~16 hours
Unsloth LoRA Llama-3-8B 2048 8 ~32,000 tok/s (1.78x) ~0.9 hours
Unsloth QLoRA Llama-3-70B 2048 1 ~3,200 tok/s (1.78×) ~9 hours

The 24GB ceiling

Method Upper limit Must-have options
LoRA fp16/bf16 ~14B Over seq 2048, use gradient checkpointing
QLoRA NF4 ~70B paged AdamW + FlashAttention-2 + double-quant required
Full-parameter SFT ~3B Optimizer CPU offload + 8-bit Adam
DPO/RLHF ~13B (LoRA) sharing reference-policy LoRAs is the key
Continued pretraining ~1~3B 8-bit Adam + grad checkpointing

Electricity bill and gut feeling

The incremental power of an 8B LoRA epoch (1.6 hours) is 350W×1.6h=0.56kWh, 120 KRW at 214.6 KRW/kWh. Even a 70B QLoRA weekend job (16 hours) costs 3,500 KRW in electricity. What’s expensive in training isn’t power but the card and time — amortizing the card (3.56M KRW/24 months) per hour is about 615 KRW/h, 5x the electricity.

The inference-side VRAM math isVRAM guide, and the runtime isRuntime comparison.

Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from sourced public measurements and my own · no local runs ·View the full hub · Quantization ladder, 104 builds · Generation-speed chart, 47 models