Qwen3.8-27B is the current best-overall local pick. The question everyone asks: does it fit 16GB, and at what speed?
Measured footprint: 17.7GB at the recommended quant — it does NOT fit a 16GB card without heavier quantization. On our 128GB M5 Max it sustains 73.2 tok/s. On 16GB you want a Q3 variant (~12GB) and expect roughly 40–55 tok/s with the quality hit that quant is known for.
Cards that run it at recommended quant today: RTX 5080 24GB-class, RTX 4090/3090 (24GB), dual 5060 Ti pooled (32GB), or any 24GB+ Mac. Korea street price for the 5080: $1,940 (16GB variant — note it will NOT fit; the 24GB/32GB SKUs are the ones that matter).
Methodology
Apple M5 Max 128GB unified memory, MLX/Ollama, tok/s = sustained generation over a 2K-token prompt, measured on this machine. Korea street prices from Danawa (multi-vendor median), converted at ~1,380 KRW/USD. Rows marked “public record” cite published benchmarks reproduced where possible; “measured” rows are from our hardware. Updated weekly — check the date in the title.
👉 Does it run on YOUR card? Check the VRAM Fit Matrix — measured estimates for every model above.

댓글 남기기