The pooled-VRAM thread that won’t die, with numbers. Bandwidth is the binding constraint for generation speed: 5070 Ti has 896 GB/s vs 448 GB/s per 5060 Ti. Pooling two 5060 Ti does not pool bandwidth for a single stream — tensor parallel roughly sums compute but adds inter-GPU overhead.
Practical read: models that fit 16GB — 5070 Ti is ~1.7x faster (published 8B Q4: 125 vs 75 tok/s). Models that don’t fit 16GB — dual 5060 Ti is the only one that runs them at all, at roughly single-5060-Ti speed plus overhead. Buy for the model size you actually run, not the spec sheet.
Korea street: 5070 Ti 16GB median $1,440, two 5060 Ti (16GB SKU prices rising — 8GB median $580). Memory supercycle warnings from Micron mean waiting has a real cost.
Methodology
Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.
👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기