GPU 시세 · VRAM · 온디바이스 AI

🌐 · English · 中文 · Español · 한국어 · · English Hub

A follow-up to yesterday’s generation-speed chart. This time it’s quantized versions of the same models. I didn’t run a single one locally. Download sizes are all measured today from Hugging Face live file listings, and quality figures are cited with explicit sources from each model card and public measurement blogs. File sizes across 104 builds for all models combined.

Conclusion first:As of 2026, 4-bit UD (Q4_K_XL class) is indistinguishable from BF16 to the human eye. Across a 300-task evaluation the deviation from BF16 stays within ±1 point. What really cuts quality isn’t quantization but the ‘thinking budget’. Even the same BF16 model scores 61.3% with thinking off and 80.3% at max — those 20 points are bigger than the entire quantization band. Only 1-bit (IQ1) genuinely collapses.

1. Format terminology — what’s inside each label

Format Stagnation Real-world evaluation
Q4_K_M llama.cpp’s default 4-bit block quantization 2024–25 industry standard. Still fine
UD-* (Unsloth Dynamic) Dynamic quantization that allocates different bits per layer. Files labeled Q2_K_XL actually mix 15 formats (the largest share is IQ3_XXS at 25.3%). Currently dominates most of the Pareto front — at the same size, go UD
IQ4_XS / IQ3_* Codebook-based low bits — small, but decoding can get slower Only when capacity is critical
MXFP4_MOE 4-bit for MoE experts only On Gemma 4, KL 0.963 loses to the same uploader’s UD-Q4_K_XL (0.747) — don’t be fooled by format names
QAT Weights retrained with quantization assumed (Google official) top-1 85.6% on Q4_0 vs 70.2% for ordinary-weights Q4_0 — best for 4-bit fixed operation
MLX nbit Apple MLX native (4/6/8bit) Mac only. Fewer files than GGUF, but a speed edge when paired with speculative decoding (mtp)
FP8 Server practice: 8-bit exponent The only range scoring above BF16 (+1.0, statistically meaningless of course)

2. Quantization ladder per model (15 models · 104 builds)

Qwen3.8-27B

Dense 27.4B · multimodal · 256K context · sourceunsloth/Qwen3.8-27B-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 54.7 GB 64GB+ Baseline 80.3%
FP8 30.9 GB 40GB+ 81.3% (+1.0, meaningless)
UD-Q8_K_XL 31.5 GB 40GB+ Effectively lossless
Q8_0 29.0 GB 32GB+ KL 0.0006 · top-1 98.9%
UD-Q6_K_XL 25.3 GB 32GB+ KL 0.0011 class
UD-Q6_K_M 23.1 GB 32GB+ 80.8% (+0.4, meaningless)
UD-Q5_K_XL 20.9 GB 24GB+ KL 0.0044-class
UD-Q5_K_M 19.8 GB 24GB+ KL 0.0042 (vs AtomicChat)
UD-Q4_K_XL 17.6 GB 24GB+ 79.6% (-0.7, meaningless) · #1 on agents, 660/720
UD-Q4_K_M 16.5 GB 24GB Agent 652 (8 points down)
Q4_0 16.1 GB 24GB No imatrix — not recommended
UD-IQ4_XS 14.3 GB 16GB HumanEval+ 86.6% (EvalPlus measurement)
UD-Q3_K_XL 13.1 GB 16GB 79.9% (-0.4) · agent 643
UD-IQ3_XXS 10.9 GB 12GB PPL 6.48 (measured on 4090)
UD-Q2_K_XL 9.8 GB 12GB 79.4% (-0.9) · 8.0-point drop in low-thinking mode
UD-IQ2_S 8.4 GB 12GB Acceptable line — not recommended for coding
UD-IQ2_XXS 7.3 GB 12GB Getting pushed out by code on 4B models
UD-IQ1_M 6.7 GB 8GB 43.0% (-37.3) — a differently compressed model

Qwen3.8-Flash-Next (125B)

MoE 125B-A95B class · 262K · measured across summed shard files · sourceunsloth/Qwen3.8-Flash-Next-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 354.1 GB 384GB+ Server-class
Q8_0 188.2 GB 192GB+ Mac Studio Ultra-class
UD-Q6_K_XL 169.2 GB 176GB+ M3 Ultra 192GB barely
UD-Q5_K_XL 158.3 GB 176GB+
UD-Q4_K_XL 111.3 GB 128GB This line starts on the M5 Max 128GB — 16GB of KV headroom
UD-IQ4_XS 93.7 GB 96GB+
UD-Q3_K_XL 90.0 GB 96GB+
UD-Q2_K_XL 78.9 GB 80GB+
UD-IQ3_XXS 82.0 GB 96GB+
UD-IQ1_M 74.5 GB 80GB+ 1-bit — quality collapse
MTP draft (Q4_K_M) 2.8 GB Additional Separate file for speculative decoding

Qwen3.6-35B-A3B

MoE 35B-A3B · 3B active · sourceunsloth/Qwen3.6-35B-A3B-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 71.9 GB 80GB+ Criterion
UD-Q8_K_XL 38.5 GB 48GB+
Q8_0 36.9 GB 48GB+
UD-Q6_K_XL 31.8 GB 40GB+
UD-Q5_K_XL 26.6 GB 32GB+
UD-Q4_K_XL 22.4 GB 24GB+ Recommended — borderline on a 4090 24GB with KV included, comfortable at 32GB
UD-Q4_K_M 22.1 GB 24GB+
UD-IQ4_XS 17.7 GB 24GB
UD-Q3_K_XL 16.8 GB 16GB+

Gemma 4 26B-A4B

MoE 26B-A4B · 128 experts · QAT co-training · sourceunsloth/gemma-4-26B-A4B-it-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 50.5 GB 64GB+ Criterion
UD-Q8_K_XL 27.6 GB 32GB+ KL 0.544 — 3.3x worse than dense peers
Q8_0 26.9 GB 32GB+ KL is already 0.54 here
UD-Q6_K_XL 23.3 GB 32GB+
UD-Q5_K_XL 21.2 GB 24GB+
UD-Q4_K_XL 17.0 GB 24GB+ KL 0.747 — the Pareto top of the 4-bit range
bartowski Q4_K_M 15.9 GB 24GB KL 1.093 — same Q4, big uploader gap
ggml/lmstudio Q4_K_M 15.9 GB 24GB KL 2.12 — avoid the Q4 from these two uploaders
MXFP4_MOE 16.6 GB 24GB KL 0.963 — even MoE-specific formats lose to UD
QAT UD-Q4_K_XL 14.2 GB 16GB+ QAT-retrained weights — Q4_0 layout
UD-Q3_K_XL 12.9 GB 16GB
UD-IQ2_XXS 9.9 GB 12GB 2-bit collapse zone

gpt-oss-20b

MoE 21B-A3.6B · weights natively MXFP4 · sourceunsloth/gpt-oss-20b-GGUF

Quantization Download Memory required Quality / notes
F16 (dense part only) 13.8 GB 16GB MoE is locked to MXFP4, so anything above this is pointless
UD-Q8_K_XL 13.2 GB 16GB Going up 1 bit costs just 0.6GB
Q8_0 12.1 GB 16GB
Q4_K_M 11.6 GB 16GB This is the practical recommended line — on a 16GB GPU, KV-inclusive is borderline; at 12GB use q8_0 KV
Q4_0 11.5 GB 12GB
Q2_K 11.5 GB 12GB Same size as Q4 — MoE just doesn’t shrink

Qwen3-Coder-30B-A3B

MoE 30B-A3B · coding-specialized · sourceunsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 61.1 GB 64GB+ Criterion
UD-Q8_K_XL 36.0 GB 48GB+
Q8_0 32.5 GB 40GB+
UD-Q6_K_XL 26.3 GB 32GB+
Q5_K_M 21.7 GB 24GB+
Q4_K_M 18.6 GB 24GB Recommended — my measured 146.2 tok/s is this tier
IQ4_XS 16.4 GB 16GB+
Q3_K_M 14.7 GB 16GB

Mistral-Small-3.2-24B

Dense 24B · sourceunsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 47.2 GB 56GB+ Criterion
UD-Q8_K_XL 29.0 GB 32GB+
Q8_0 25.1 GB 32GB+
UD-Q6_K_XL 20.8 GB 24GB+
UD-Q4_K_XL 14.5 GB 16GB+ Recommended
IQ4_XS 12.8 GB 16GB

DeepSeek-R1-Distill-8B

Dense 8B distillation · sourceunsloth/DeepSeek-R1-Distill-Llama-8B-GGUF

Quantization Download Memory required Quality / notes
F16 (original) 16.1 GB 24GB Criterion
UD-Q8_K_XL 10.6 GB 12GB
Q8_0 8.5 GB 12GB
UD-Q4_K_XL 5.0 GB 8GB Recommended
Q4_K_M 4.9 GB 8GB

Phi-4

Dense 14.6B · Sourceunsloth/phi-4-GGUF

Quantization Download Memory required Quality / notes
F16 (original) 29.3 GB 32GB+ Criterion
Q8_0 15.6 GB 24GB
Q6_K 12.0 GB 16GB
Q5_K_M 10.4 GB 16GB
Q4_K_M 8.9 GB 12GB Recommended — comfortable on a 4090 including KV
Q3_K_M 7.2 GB 8GB
Q2_K 5.6 GB 8GB Being a reasoning model, its reasoning breaks first at 2-bit

Llama-3.1-8B

Dense 8B · sourcebartowski/Meta-Llama-3.1-8B-Instruct-GGUF

Quantization Download Memory required Quality / notes
Q8_0 8.5 GB 12GB
Q6_K 6.6 GB 8GB
Q5_K_M 5.7 GB 8GB
Q4_K_M 4.9 GB 8GB Recommended — effectively the industry default
IQ4_XS 4.4 GB 6GB
Q3_K_M 4.0 GB 6GB

Ornith-1.5-35B-A3B

MoE 35B-A3B · new (26-08) · sourceornith-ai/Ornith-1.5-35B-A3B-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 71.1 GB 80GB+ Criterion
Q8_0 37.8 GB 48GB+
Q6_K 29.2 GB 32GB+
Q5_K_M 25.3 GB 32GB+
Q4_K_M 21.7 GB 24GB Official single ladder — my public benchmark record is 8bit MLX (92.6 tok/s, M5 Max)

Nemotron-3.5-Lightning-30B-A3B

MoE 30B-A3B · built-in MTP head · sourceggml-org/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 63.2 GB 80GB+ Criterion
Q8_0 33.6 GB 40GB+
Q4_0 18.9 GB 24GB Only the 3 official ggml ones — my public-record bench is MLX 4bit (114.9 tok/s, M5 Max)
MTP draft (Q4_0) 1.2 GB Additional Speculative decoding heads loaded separately

Muse-Glimmer-30B

Dense 30B · new from Meta (26-08) · sourcemeta-models/Muse-Glimmer-30B-GGUF

Quantization Download Memory required Quality / notes
Q4_K_M (17GB) 16.8 GB 24GB Meta official — my benchmark’s 26.6 tok/s is this class (M5 Max)
Q4_K_XL (dynamic) 19.7 GB 24GB Only 2 official variants
dflash draft 1.6 GB Additional For speculative decoding — the 33x acceleration card is based on this combo

Laguna-XS-2.1

MoE 34B-class · Poolside new (26-06) · Sourceggml-org/Laguna-XS-2.1-GGUF

Quantization Download Memory required Quality / notes
BF16 (original) 66.9 GB 80GB+ Criterion
Q8_0 35.6 GB 48GB+
Q4_K_M 19.6 GB 24GB Recommended — my public benchmark record is MLX 4bit (126.0 tok/s, M5 Max)
APEX balanced (mudler) 24.3 GB 32GB Gate report — three graded tiers (Quality/Balanced/Compact)

3. Where quality actually breaks

Quantization quality — where and how much it breaks: measured case of Qwen3.8-27B

One measurement published the BF16 comparison using 300 fixed tasks (pass@1) and a 240-task agent ladder (720 points max):

Build Download pass@1 (vs BF16 80.3%) Agent /720
BF16 54.7 GB 80.3% (baseline) Not measured
FP8 30.9 GB 81.3% (+1.0, p=0.49) 647
UD-Q6_K_M 23.1 GB 80.8% (+0.4, p=0.76) Not measured
UD-Q4_K_XL 17.6 GB 79.6% (-0.7, p=0.58) 660 (1st across all bands)
UD-Q4_K_M 16.5 GB Not measured 652
UD-Q3_K_XL 13.1 GB 79.9% (-0.4, p=0.79) 643
UD-Q2_K_XL 9.8 GB 79.4% (-0.9, p=0.54) 629
UD-IQ1_M 6.7 GB 43.0% (-37.3, p<0.0001) Not measured

The most important line here isn’t the last one — it’s Q4_K_XL. The 17.6GB build, compared to BF1613 points higher on agent tasks. Statistically a tie, yet first on the agent ladder — quantization noise appears to actually diversify reasoning paths, though the benchmarkers summarized it as ‘no measurable degradation’. Back to the reasoning-budget point: the same model in BF16, reasoning off 61.3% → max 80.3%. That’s 7 times the entire quantization band (±3 points). Unless you’re on Q4 because GPU VRAM is too small, spending the saved VRAM on longer context and thinking wins.

MoE dislikes quantization even more — KL measurements across 80 variants on Gemma 4 26B-A4B

There are measurements of KL divergence against BF16 across 80 builds from 6 uploaders, using about 250K tokens (coding, chat, tool calling, science, non-Latin scripts, long context). Summary:

  • A 4B-active MoE is already at KL 0.544 on Q8_0 — 3.3x that of a dense 31B peer (0.163). MoE is simply weak to quantization.
  • Same Q4_K_M, yet ggml-org 2.126, lmstudio 2.116, bartowski 1.093 —The uploader matters more than the quantization name.
  • unsloth UD monopolizes the Pareto front (one exception: bartowski Q6_K_L).
  • The MoE-only format MXFP4_MOE (16.6GB, KL 0.963) loses to the same family’s UD-Q4_K_XL (17.0GB, KL 0.747).

And there’s a measurement showing calibration data decides real task performance: in the same recipe, switching only the calibration data from code/math to mixed improved KL but dropped GSM8K by another 9%P. It’s a public proof that ‘KL improvement = task quality improvement’ does not hold.

4. Recommended matrix for my machine

Model M5 Max 128GB recommended RTX 4090 24GB recommended Evidence
Qwen3.8-27B UD-Q6_K_XL (25.3) UD-Q4_K_XL (17.6) Degradation meaningless from 4-bit on; headroom with KV included at 24GB
Qwen3.8-Flash-Next UD-Q4_K_XL (111.3) Exceeds specs Fits in 128GB with 16GB left for KV
Qwen3.6-35B-A3B Q8_0 (36.9) UD-Q4_K_M (22.1)+q8 KV Fast even at 4-bit thanks to 3B active
Gemma 4 26B-A4B UD-Q6_K_XL (23.3) QAT UD-Q4_K_XL (14.2) KL-benchmarked Pareto — even 16GB GPUs fit with QAT
gpt-oss-20b Q8_0 (12.1) Q4_K_M (11.6) MoE is natively MXFP4 — the ladder means little
Qwen3-Coder-30B Q8_0 (32.5) Q4_K_M (18.6) My measured 146.2 tok/s range
Ornith-1.5-35B Q8_0 (37.8) Q4_K_M (21.7) Official ladder, 5 tiers
Nemotron-3.5-Lightning Q8_0 (33.6) Q4_0 (18.9) 3 official variants + MTP head
Muse-Glimmer-30B Q4_K_XL (19.7) Q4_K_M (16.8) Only 2 official variants
Laguna-XS-2.1 Q8_0 (35.6) Q4_K_M (19.6) ggml formula
Phi-4 Q8_0 (15.6) Q4_K_M (8.9) Lightweight utility
Llama-3.1-8B Q8_0 (8.5) Q4_K_M (4.9) Momentum maintained

Size vs memory — a formula you memorize by measurement

The real RAM you need on top of the download size looks roughly like this: weights + (for a 16K context) 1~4GB KV cache (halved with a q8_0 cache) + 1~2GB overhead. Example: Qwen3.8-27B Q4_K_XL 17.6GB + 16K context ≈ 20GB → fits on a 24GB card. With M5 Max unified memory this term grows the longer you push the context, so even a 128GB machine can’t handle a 262K context on builds over 110GB (Flash-Next Q5 and up). The reason Flash-Next MLX held 59 tok/s in yesterday’s chart is the 4bit + SSD offload combo; anything above that bit depth is a luxury on this machine.

5. Sources and Cautions

Source

Caution: PPL/KL across different harnesses cannot be compared in absolute terms — only comparisons within the same table are valid (the measurers’ common disclaimer).

This pageLocal LLM Hub — What Runs on My Computeris part of. The per-hardware supported-model table is in the hub.