🌐 · English · 中文 · Español · 한국어 · · English Hub
A follow-up to yesterday’s generation-speed chart. This time it’s quantized versions of the same models. I didn’t run a single one locally. Download sizes are all measured today from Hugging Face live file listings, and quality figures are cited with explicit sources from each model card and public measurement blogs. File sizes across 104 builds for all models combined.
1. Format terminology — what’s inside each label
| Format | Stagnation | Real-world evaluation |
|---|---|---|
| Q4_K_M | llama.cpp’s default 4-bit block quantization | 2024–25 industry standard. Still fine |
| UD-* (Unsloth Dynamic) | Dynamic quantization that allocates different bits per layer. Files labeled Q2_K_XL actually mix 15 formats (the largest share is IQ3_XXS at 25.3%). | Currently dominates most of the Pareto front — at the same size, go UD |
| IQ4_XS / IQ3_* | Codebook-based low bits — small, but decoding can get slower | Only when capacity is critical |
| MXFP4_MOE | 4-bit for MoE experts only | On Gemma 4, KL 0.963 loses to the same uploader’s UD-Q4_K_XL (0.747) — don’t be fooled by format names |
| QAT | Weights retrained with quantization assumed (Google official) | top-1 85.6% on Q4_0 vs 70.2% for ordinary-weights Q4_0 — best for 4-bit fixed operation |
| MLX nbit | Apple MLX native (4/6/8bit) | Mac only. Fewer files than GGUF, but a speed edge when paired with speculative decoding (mtp) |
| FP8 | Server practice: 8-bit exponent | The only range scoring above BF16 (+1.0, statistically meaningless of course) |
2. Quantization ladder per model (15 models · 104 builds)
Qwen3.8-27B
Dense 27.4B · multimodal · 256K context · sourceunsloth/Qwen3.8-27B-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 54.7 GB | 64GB+ | Baseline 80.3% |
| FP8 | 30.9 GB | 40GB+ | 81.3% (+1.0, meaningless) |
| UD-Q8_K_XL | 31.5 GB | 40GB+ | Effectively lossless |
| Q8_0 | 29.0 GB | 32GB+ | KL 0.0006 · top-1 98.9% |
| UD-Q6_K_XL | 25.3 GB | 32GB+ | KL 0.0011 class |
| UD-Q6_K_M | 23.1 GB | 32GB+ | 80.8% (+0.4, meaningless) |
| UD-Q5_K_XL | 20.9 GB | 24GB+ | KL 0.0044-class |
| UD-Q5_K_M | 19.8 GB | 24GB+ | KL 0.0042 (vs AtomicChat) |
| UD-Q4_K_XL | 17.6 GB | 24GB+ | 79.6% (-0.7, meaningless) · #1 on agents, 660/720 |
| UD-Q4_K_M | 16.5 GB | 24GB | Agent 652 (8 points down) |
| Q4_0 | 16.1 GB | 24GB | No imatrix — not recommended |
| UD-IQ4_XS | 14.3 GB | 16GB | HumanEval+ 86.6% (EvalPlus measurement) |
| UD-Q3_K_XL | 13.1 GB | 16GB | 79.9% (-0.4) · agent 643 |
| UD-IQ3_XXS | 10.9 GB | 12GB | PPL 6.48 (measured on 4090) |
| UD-Q2_K_XL | 9.8 GB | 12GB | 79.4% (-0.9) · 8.0-point drop in low-thinking mode |
| UD-IQ2_S | 8.4 GB | 12GB | Acceptable line — not recommended for coding |
| UD-IQ2_XXS | 7.3 GB | 12GB | Getting pushed out by code on 4B models |
| UD-IQ1_M | 6.7 GB | 8GB | 43.0% (-37.3) — a differently compressed model |
Qwen3.8-Flash-Next (125B)
MoE 125B-A95B class · 262K · measured across summed shard files · sourceunsloth/Qwen3.8-Flash-Next-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 354.1 GB | 384GB+ | Server-class |
| Q8_0 | 188.2 GB | 192GB+ | Mac Studio Ultra-class |
| UD-Q6_K_XL | 169.2 GB | 176GB+ | M3 Ultra 192GB barely |
| UD-Q5_K_XL | 158.3 GB | 176GB+ | |
| UD-Q4_K_XL | 111.3 GB | 128GB | This line starts on the M5 Max 128GB — 16GB of KV headroom |
| UD-IQ4_XS | 93.7 GB | 96GB+ | |
| UD-Q3_K_XL | 90.0 GB | 96GB+ | |
| UD-Q2_K_XL | 78.9 GB | 80GB+ | |
| UD-IQ3_XXS | 82.0 GB | 96GB+ | |
| UD-IQ1_M | 74.5 GB | 80GB+ | 1-bit — quality collapse |
| MTP draft (Q4_K_M) | 2.8 GB | Additional | Separate file for speculative decoding |
Qwen3.6-35B-A3B
MoE 35B-A3B · 3B active · sourceunsloth/Qwen3.6-35B-A3B-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 71.9 GB | 80GB+ | Criterion |
| UD-Q8_K_XL | 38.5 GB | 48GB+ | |
| Q8_0 | 36.9 GB | 48GB+ | |
| UD-Q6_K_XL | 31.8 GB | 40GB+ | |
| UD-Q5_K_XL | 26.6 GB | 32GB+ | |
| UD-Q4_K_XL | 22.4 GB | 24GB+ | Recommended — borderline on a 4090 24GB with KV included, comfortable at 32GB |
| UD-Q4_K_M | 22.1 GB | 24GB+ | |
| UD-IQ4_XS | 17.7 GB | 24GB | |
| UD-Q3_K_XL | 16.8 GB | 16GB+ |
Gemma 4 26B-A4B
MoE 26B-A4B · 128 experts · QAT co-training · sourceunsloth/gemma-4-26B-A4B-it-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 50.5 GB | 64GB+ | Criterion |
| UD-Q8_K_XL | 27.6 GB | 32GB+ | KL 0.544 — 3.3x worse than dense peers |
| Q8_0 | 26.9 GB | 32GB+ | KL is already 0.54 here |
| UD-Q6_K_XL | 23.3 GB | 32GB+ | |
| UD-Q5_K_XL | 21.2 GB | 24GB+ | |
| UD-Q4_K_XL | 17.0 GB | 24GB+ | KL 0.747 — the Pareto top of the 4-bit range |
| bartowski Q4_K_M | 15.9 GB | 24GB | KL 1.093 — same Q4, big uploader gap |
| ggml/lmstudio Q4_K_M | 15.9 GB | 24GB | KL 2.12 — avoid the Q4 from these two uploaders |
| MXFP4_MOE | 16.6 GB | 24GB | KL 0.963 — even MoE-specific formats lose to UD |
| QAT UD-Q4_K_XL | 14.2 GB | 16GB+ | QAT-retrained weights — Q4_0 layout |
| UD-Q3_K_XL | 12.9 GB | 16GB | |
| UD-IQ2_XXS | 9.9 GB | 12GB | 2-bit collapse zone |
gpt-oss-20b
MoE 21B-A3.6B · weights natively MXFP4 · sourceunsloth/gpt-oss-20b-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| F16 (dense part only) | 13.8 GB | 16GB | MoE is locked to MXFP4, so anything above this is pointless |
| UD-Q8_K_XL | 13.2 GB | 16GB | Going up 1 bit costs just 0.6GB |
| Q8_0 | 12.1 GB | 16GB | |
| Q4_K_M | 11.6 GB | 16GB | This is the practical recommended line — on a 16GB GPU, KV-inclusive is borderline; at 12GB use q8_0 KV |
| Q4_0 | 11.5 GB | 12GB | |
| Q2_K | 11.5 GB | 12GB | Same size as Q4 — MoE just doesn’t shrink |
Qwen3-Coder-30B-A3B
MoE 30B-A3B · coding-specialized · sourceunsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 61.1 GB | 64GB+ | Criterion |
| UD-Q8_K_XL | 36.0 GB | 48GB+ | |
| Q8_0 | 32.5 GB | 40GB+ | |
| UD-Q6_K_XL | 26.3 GB | 32GB+ | |
| Q5_K_M | 21.7 GB | 24GB+ | |
| Q4_K_M | 18.6 GB | 24GB | Recommended — my measured 146.2 tok/s is this tier |
| IQ4_XS | 16.4 GB | 16GB+ | |
| Q3_K_M | 14.7 GB | 16GB |
Mistral-Small-3.2-24B
Dense 24B · sourceunsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 47.2 GB | 56GB+ | Criterion |
| UD-Q8_K_XL | 29.0 GB | 32GB+ | |
| Q8_0 | 25.1 GB | 32GB+ | |
| UD-Q6_K_XL | 20.8 GB | 24GB+ | |
| UD-Q4_K_XL | 14.5 GB | 16GB+ | Recommended |
| IQ4_XS | 12.8 GB | 16GB |
DeepSeek-R1-Distill-8B
Dense 8B distillation · sourceunsloth/DeepSeek-R1-Distill-Llama-8B-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| F16 (original) | 16.1 GB | 24GB | Criterion |
| UD-Q8_K_XL | 10.6 GB | 12GB | |
| Q8_0 | 8.5 GB | 12GB | |
| UD-Q4_K_XL | 5.0 GB | 8GB | Recommended |
| Q4_K_M | 4.9 GB | 8GB |
Phi-4
Dense 14.6B · Sourceunsloth/phi-4-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| F16 (original) | 29.3 GB | 32GB+ | Criterion |
| Q8_0 | 15.6 GB | 24GB | |
| Q6_K | 12.0 GB | 16GB | |
| Q5_K_M | 10.4 GB | 16GB | |
| Q4_K_M | 8.9 GB | 12GB | Recommended — comfortable on a 4090 including KV |
| Q3_K_M | 7.2 GB | 8GB | |
| Q2_K | 5.6 GB | 8GB | Being a reasoning model, its reasoning breaks first at 2-bit |
Llama-3.1-8B
Dense 8B · sourcebartowski/Meta-Llama-3.1-8B-Instruct-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| Q8_0 | 8.5 GB | 12GB | |
| Q6_K | 6.6 GB | 8GB | |
| Q5_K_M | 5.7 GB | 8GB | |
| Q4_K_M | 4.9 GB | 8GB | Recommended — effectively the industry default |
| IQ4_XS | 4.4 GB | 6GB | |
| Q3_K_M | 4.0 GB | 6GB |
Ornith-1.5-35B-A3B
MoE 35B-A3B · new (26-08) · sourceornith-ai/Ornith-1.5-35B-A3B-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 71.1 GB | 80GB+ | Criterion |
| Q8_0 | 37.8 GB | 48GB+ | |
| Q6_K | 29.2 GB | 32GB+ | |
| Q5_K_M | 25.3 GB | 32GB+ | |
| Q4_K_M | 21.7 GB | 24GB | Official single ladder — my public benchmark record is 8bit MLX (92.6 tok/s, M5 Max) |
Nemotron-3.5-Lightning-30B-A3B
MoE 30B-A3B · built-in MTP head · sourceggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 63.2 GB | 80GB+ | Criterion |
| Q8_0 | 33.6 GB | 40GB+ | |
| Q4_0 | 18.9 GB | 24GB | Only the 3 official ggml ones — my public-record bench is MLX 4bit (114.9 tok/s, M5 Max) |
| MTP draft (Q4_0) | 1.2 GB | Additional | Speculative decoding heads loaded separately |
Muse-Glimmer-30B
Dense 30B · new from Meta (26-08) · sourcemeta-models/Muse-Glimmer-30B-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| Q4_K_M (17GB) | 16.8 GB | 24GB | Meta official — my benchmark’s 26.6 tok/s is this class (M5 Max) |
| Q4_K_XL (dynamic) | 19.7 GB | 24GB | Only 2 official variants |
| dflash draft | 1.6 GB | Additional | For speculative decoding — the 33x acceleration card is based on this combo |
Laguna-XS-2.1
MoE 34B-class · Poolside new (26-06) · Sourceggml-org/Laguna-XS-2.1-GGUF
| Quantization | Download | Memory required | Quality / notes |
|---|---|---|---|
| BF16 (original) | 66.9 GB | 80GB+ | Criterion |
| Q8_0 | 35.6 GB | 48GB+ | |
| Q4_K_M | 19.6 GB | 24GB | Recommended — my public benchmark record is MLX 4bit (126.0 tok/s, M5 Max) |
| APEX balanced (mudler) | 24.3 GB | 32GB | Gate report — three graded tiers (Quality/Balanced/Compact) |
3. Where quality actually breaks
Quantization quality — where and how much it breaks: measured case of Qwen3.8-27B
One measurement published the BF16 comparison using 300 fixed tasks (pass@1) and a 240-task agent ladder (720 points max):
| Build | Download | pass@1 (vs BF16 80.3%) | Agent /720 |
|---|---|---|---|
| BF16 | 54.7 GB | 80.3% (baseline) | Not measured |
| FP8 | 30.9 GB | 81.3% (+1.0, p=0.49) | 647 |
| UD-Q6_K_M | 23.1 GB | 80.8% (+0.4, p=0.76) | Not measured |
| UD-Q4_K_XL | 17.6 GB | 79.6% (-0.7, p=0.58) | 660 (1st across all bands) |
| UD-Q4_K_M | 16.5 GB | Not measured | 652 |
| UD-Q3_K_XL | 13.1 GB | 79.9% (-0.4, p=0.79) | 643 |
| UD-Q2_K_XL | 9.8 GB | 79.4% (-0.9, p=0.54) | 629 |
| UD-IQ1_M | 6.7 GB | 43.0% (-37.3, p<0.0001) | Not measured |
The most important line here isn’t the last one — it’s Q4_K_XL. The 17.6GB build, compared to BF1613 points higher on agent tasks. Statistically a tie, yet first on the agent ladder — quantization noise appears to actually diversify reasoning paths, though the benchmarkers summarized it as ‘no measurable degradation’. Back to the reasoning-budget point: the same model in BF16, reasoning off 61.3% → max 80.3%. That’s 7 times the entire quantization band (±3 points). Unless you’re on Q4 because GPU VRAM is too small, spending the saved VRAM on longer context and thinking wins.
MoE dislikes quantization even more — KL measurements across 80 variants on Gemma 4 26B-A4B
There are measurements of KL divergence against BF16 across 80 builds from 6 uploaders, using about 250K tokens (coding, chat, tool calling, science, non-Latin scripts, long context). Summary:
- A 4B-active MoE is already at KL 0.544 on Q8_0 — 3.3x that of a dense 31B peer (0.163). MoE is simply weak to quantization.
- Same Q4_K_M, yet ggml-org 2.126, lmstudio 2.116, bartowski 1.093 —The uploader matters more than the quantization name.
- unsloth UD monopolizes the Pareto front (one exception: bartowski Q6_K_L).
- The MoE-only format MXFP4_MOE (16.6GB, KL 0.963) loses to the same family’s UD-Q4_K_XL (17.0GB, KL 0.747).
And there’s a measurement showing calibration data decides real task performance: in the same recipe, switching only the calibration data from code/math to mixed improved KL but dropped GSM8K by another 9%P. It’s a public proof that ‘KL improvement = task quality improvement’ does not hold.
4. Recommended matrix for my machine
| Model | M5 Max 128GB recommended | RTX 4090 24GB recommended | Evidence |
|---|---|---|---|
| Qwen3.8-27B | UD-Q6_K_XL (25.3) | UD-Q4_K_XL (17.6) | Degradation meaningless from 4-bit on; headroom with KV included at 24GB |
| Qwen3.8-Flash-Next | UD-Q4_K_XL (111.3) | Exceeds specs | Fits in 128GB with 16GB left for KV |
| Qwen3.6-35B-A3B | Q8_0 (36.9) | UD-Q4_K_M (22.1)+q8 KV | Fast even at 4-bit thanks to 3B active |
| Gemma 4 26B-A4B | UD-Q6_K_XL (23.3) | QAT UD-Q4_K_XL (14.2) | KL-benchmarked Pareto — even 16GB GPUs fit with QAT |
| gpt-oss-20b | Q8_0 (12.1) | Q4_K_M (11.6) | MoE is natively MXFP4 — the ladder means little |
| Qwen3-Coder-30B | Q8_0 (32.5) | Q4_K_M (18.6) | My measured 146.2 tok/s range |
| Ornith-1.5-35B | Q8_0 (37.8) | Q4_K_M (21.7) | Official ladder, 5 tiers |
| Nemotron-3.5-Lightning | Q8_0 (33.6) | Q4_0 (18.9) | 3 official variants + MTP head |
| Muse-Glimmer-30B | Q4_K_XL (19.7) | Q4_K_M (16.8) | Only 2 official variants |
| Laguna-XS-2.1 | Q8_0 (35.6) | Q4_K_M (19.6) | ggml formula |
| Phi-4 | Q8_0 (15.6) | Q4_K_M (8.9) | Lightweight utility |
| Llama-3.1-8B | Q8_0 (8.5) | Q4_K_M (4.9) | Momentum maintained |
Size vs memory — a formula you memorize by measurement
The real RAM you need on top of the download size looks roughly like this: weights + (for a 16K context) 1~4GB KV cache (halved with a q8_0 cache) + 1~2GB overhead. Example: Qwen3.8-27B Q4_K_XL 17.6GB + 16K context ≈ 20GB → fits on a 24GB card. With M5 Max unified memory this term grows the longer you push the context, so even a 128GB machine can’t handle a 262K context on builds over 110GB (Flash-Next Q5 and up). The reason Flash-Next MLX held 59 tok/s in yesterday’s chart is the 4bit + SSD offload combo; anything above that bit depth is a luxury on this machine.
5. Sources and Cautions
Source
- All file sizes:
unsloth/Qwen3.8-27B-GGUFList of 14 additional repo files (measured via API 2026-09-27, shards summed) - Qwen3.8-27B pass@1 / agent ladder: vramcalculator.com quantization report (@superalesha’s fixed 300-task measurement is the original source)
- Qwen3.8-27B KL/top-1 table:
AtomicChat/Qwen3.8-27B-GGUF(vs BF16 reference, 87 held-out chunks) - 80 Gemma 4 KL variants: localbench.substack.com GGUF Quality Benchmark
- Gemma 4 QAT top-1:
SC117/gemma-4-26B-A4B-it-qat-heretic-GGUFModel card - HumanEval+/EvalPlus table:
tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF - Unleashed Q3 vs Q4 PPL/token speed:
outsourc-e/Qwen3.8-27B-Unleashed-GGUF(4090, llama.cpp)
Caution: PPL/KL across different harnesses cannot be compared in absolute terms — only comparisons within the same table are valid (the measurers’ common disclaimer).
This pageLocal LLM Hub — What Runs on My Computeris part of. The per-hardware supported-model table is in the hub.