What you can train on a single 4090 24GB, and at what cost — the recipe table is gigagpu’s measured data, with Qwen mappings added to suit my machine.
🌐 · English · 中文 · Español · 한국어 · · English Hub
| Recipes | Model | Sequence | Batch | Processing speed | 50k samples, 1 epoch |
|---|---|---|---|---|---|
| LoRA bf16 | Llama-3-8B (equivalent to Qwen3-8B) | 2048 | 8 | ~18,000 tok/s | ~1.6 hours |
| LoRA bf16 | Mistral-7B | 2048 | 8 | ~21,000 tok/s | ~1.4 hours |
| LoRA bf16 | Qwen2.5-14B | 2048 | 2 | ~6,400 tok/s | ~4.4 hours |
| QLoRA NF4 | Llama-3-8B | 2048 | 8 | ~14,500 tok/s | ~2.0 hours |
| QLoRA NF4 | Llama-3-70B | 2048 | 1 | ~1,800 tok/s | ~16 hours |
| Unsloth LoRA | Llama-3-8B | 2048 | 8 | ~32,000 tok/s (1.78x) | ~0.9 hours |
| Unsloth QLoRA | Llama-3-70B | 2048 | 1 | ~3,200 tok/s (1.78×) | ~9 hours |
The 24GB ceiling
| Method | Upper limit | Must-have options |
|---|---|---|
| LoRA fp16/bf16 | ~14B | Over seq 2048, use gradient checkpointing |
| QLoRA NF4 | ~70B | paged AdamW + FlashAttention-2 + double-quant required |
| Full-parameter SFT | ~3B | Optimizer CPU offload + 8-bit Adam |
| DPO/RLHF | ~13B (LoRA) | sharing reference-policy LoRAs is the key |
| Continued pretraining | ~1~3B | 8-bit Adam + grad checkpointing |
Electricity bill and gut feeling
The incremental power of an 8B LoRA epoch (1.6 hours) is 350W×1.6h=0.56kWh, 120 KRW at 214.6 KRW/kWh. Even a 70B QLoRA weekend job (16 hours) costs 3,500 KRW in electricity. What’s expensive in training isn’t power but the card and time — amortizing the card (3.56M KRW/24 months) per hour is about 615 KRW/h, 5x the electricity.
The inference-side VRAM math isVRAM guide, and the runtime isRuntime comparison.
Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from sourced public measurements and my own · no local runs ·View the full hub · Quantization ladder, 104 builds · Generation-speed chart, 47 models