GPU 시세 · VRAM · 온디바이스 AI

2026 Local LLM Benchmarks: Full Comparison Table of 47 Models

Updated:Generated September 27, 2026. Benchmarks, quantizations, and memory requirements for 47 major local LLMs compared in one shot.
Source:HuggingFace Leaderboard, LMArena Arena ELO, Benchmarklist.com ECI, official model technical reports. All values are public benchmark scores.
Updated:Scheduled to update on the 1st of each month.Daily Deep Learningblog.

🌐 · English · 中文 · Español · 한국어 · · English Hub

🏆 Section 1: One-Stop Comparison of Top Models (ECI + Arena)

Top 10 models as of September 2026, ranked by benchmark ECI and Arena ELO scores.

Rank Model name Parameters Architecture Vendor ECI Arena ELO GPQA SWE-bench HuggingFace
1 Qwen3.8 Max 397B Dense Qwen 147.85 1,481 92.7% 67.7%
2 Qwen3.8 27B 27B Dense Qwen 141.90 1,200 90.5% 58.0%
3 Gemma 4 31B 31B Dense Google 128.20 — — —
4 Mistral Large 3 128B 128B Dense Mistral 123.00 — — —
5 DeepSeek-V3.1 671B (37B active) MoE DeepSeek 122.80 — — —
6 MiniMax-M3 428B Dense MiniMax 120.00 — — —
7 Kimi K3 1.1T MoE Moonshot 120.00 — — —
8 Qwen3.8 Flash-Next 125B (6B active) MoE Qwen 120.00 1,100 86.0% —
9 Command-A+ 218B Sparse Cohere 118.00 — — —
10 Mixtral 8x22B 141B (39B active) MoE Mistral 95.00 — — —
ECI (Estimated Context Index):Composite benchmark score provided by benchmarklist.com. Computed by aggregating the model’s text/coding/math/creativity/expert-knowledge ability.
Arena ELO:Model competitiveness ranking computed from user votes on LMArena (LMSYS).
GPQA:GPQA Diamond benchmark — graduate-level science problem-solving ability.SWE-bench:Software engineering benchmark — real code-fixing ability.

📊 Section 2: Full 47-Model Comparison Table

Benchmarks, architecture, and public info for every local LLM, sorted by vendor.

Qwen series (Qwen team)

Model name Parameters Context Release ECI Arena Rk GPQA SWE-bench D/L HuggingFace
Qwen3.8 Max 397B 262K Aug 2026 147.85 21 92.7% 67.7% 50K
Qwen3.8 27B 27B 262K Aug 2026 141.90 88 90.5% 58.0% 6.7M
Qwen3.8 Flash-Next 125B (6B active) 256K Aug 2026 120.00 149 86.0% — 1.2M
Qwen3.6 35B-A3B 35B (3B active) 262K Jul 2026 125.70 109 85.0% — 3.3M
Qwen3.6 27B 27B 262K Jul 2026 118.00 61 89.0% 50.0% 2.7M
Qwen3.5 397B-A17B 397B (17B active) 262K Jun 2026 108.00 81 88.0% — 305K
Qwen3-32B 32B 128K Mar 2025 110.00 208 — — 4.0M
Qwen3-235B-A22B 235B (22B active) 256K Mar 2025 97.00 177 — — 371K
Qwen3-VL 235B-A22B 235B (22B active) 256K May 2026 96.00 153 — — 326K
Qwen2.5-72B Instruct 72B 128K Nov 2025 82.00 179 83.0% — 294K
Qwen2.5-Coder 32B 32B 128K Nov 2025 105.00 166 84.0% — 1.0M
QwQ-32B 32B 32K Jan 2025 108.00 108 — — 70.6K

Meta Llama series

Model name Parameters Context Release ECI Arena Rk GPQA SWE-bench D/L HuggingFace
Llama 4 Maverick 17B-128E 17B (128 experts) 128K Sep 2025 110.00 — — — 8.2K
Llama 4 Scout 17B-16E 17B (16 experts) 128K Sep 2025 115.00 — — — 26.2K
Llama 3.3 70B 70B 128K Sep 2025 112.00 — — — —
Llama 3.1 70B 70B 128K Jul 2025 108.00 — — — —
Llama 3.1 8B 8B 128K Jul 2025 92.00 — — — —
Llama 3.2 3B 3B 128K May 2024 75.00 — — — —

Google Gemma series

Model name Parameters Context Release ECI Arena Rk GPQA SWE-bench D/L HuggingFace
Gemma 4 31B 31B 8K Jun 2025 128.20 — — — —
Gemma 4 26B-A4B 26B (4B active) 8K Jun 2025 118.00 — — — —
Gemma 3 27B 27B 128K Apr 2025 113.00 — — — —
Gemma 3 12B 12B 128K Apr 2025 105.00 — — — —

Mistral series

Model name Parameters Context Release ECI Arena Rk GPQA SWE-bench D/L HuggingFace
Mistral Large 3 128B 128B 128K Jul 2025 123.00 — — — —
Mistral Small 3.2 24B 24B 128K Jul 2025 115.00 — — — —
Mixtral 8x22B Instruct 141B (39B active) 32K Mar 2024 95.00 — — — —
Mixtral 8x7B Instruct 47B (12B active) 32K Dec 2023 85.00 — — — —

DeepSeek series

Model name Parameters Context Release ECI Arena Rk GPQA SWE-bench D/L HuggingFace
DeepSeek-V3.1 671B (37B active) 128K Jan 2025 122.80 — — — —
DeepSeek-V3.2 685B (37B active) 128K Jan 2025 115.00 — — — —
DeepSeek-R1 671B (37B active) 128K Dec 2024 113.00 — — — —
DeepSeek-V3 671B (37B active) 128K Dec 2024 112.00 — — — —
DeepSeek-Coder-V2 16B (27B active) 128K May 2024 105.00 — — — —

Other major models

Model name Parameters Context Release ECI Arena Rk GPQA SWE-bench D/L HuggingFace
GLM-4.7 (Zhipu AI) 355B 128K Jul 2025 115.00 — — — —
GLM-4 Plus (Zhipu AI) 176B 128K May 2025 110.00 — — — —
Command-A+ (Cohere) 218B 128K Aug 2025 118.00 — — — —
Command-R+ (Cohere) 104B 128K Sep 2024 108.00 — — — —
Phi-4 (Microsoft) 14B 16K Mar 2025 98.00 — — — —
Phi-3 Medium (Microsoft) 14B 4K Mar 2024 82.00 — — — —
Nemotron-4 340B (NVIDIA) 340B 128K Jun 2025 108.00 — — — —
Nemotron-3 Super 120B (NVIDIA) 120B (12B active) 128K Jun 2025 105.00 — — — —
OLMo-3 32B Think (Allen AI) 32B 128K Jul 2025 96.00 — — — —
OLMo-3 32B (Allen AI) 32B 128K Jul 2025 94.00 — — — —
Step-3 (StepFun) 198B 128K Aug 2025 108.00 — — — —
Kimi K3 (Moonshot) 1.1T 128K Jul 2025 120.00 — — — —
MiMo-V2.5 Pro (Xiaomi) 1.02T 128K Aug 2025 115.00 — — — —
InternLM 2.5 20B (InternLM) 20B 128K Aug 2024 92.00 — — — —
Solar 10.7B (Upstage) 10.7B 128K Sep 2024 90.00 — — — —

⚡ Section 3: Quantization Comparison Table

Available quantization formats and memory requirements per model.

Qwen3.8 27B quantization variants

Form Parameter count Memory Download URL
safetensors (Original) 27.4B 55.8 GB 2.82M
Q4_K_M 27.4B 18.4 GB —
Q4_K_S 27.4B 15.7 GB —
Q8_0 27.4B 30.5 GB —
Q5_1 27.4B 22.4 GB —
Q5_K_M 27.4B 21.1 GB —
Q3_K_L 27.4B 14.5 GB —
Q3_K_M 27.4B 13.2 GB —
Q2_K 27.4B 10.5 GB —
Memory reference:For a 27B model, Q4_K_M (18.4GB) → runs on RTX 4090 (24GB). Q8_0 (30.5GB) → needs RTX 4090 SLI or an A100. 4-bit quantization is the most efficient for local execution.

🎯 Section 4: Key Achievements Summary

Qwen3.8 Max

GPQA Diamond:92.7% — top score (GPT-4o: 88.9%, Claude Opus 4.6: 85.0%)
SWE-bench Verified:67.7% — top score (GPT-4.1: 57.2%, Claude Opus 4.6: 54.5%)
LiveCodeBench:93.8% — top score (GPT-4.1: 89.4%)
SWE-bench Full:63.0% — top score
MMLU-Pro:87.8% — highest score (GPT-4.1: 85.6%)
HumanEval:91.8% — top score (GPT-4.1: 90.4%)
GPQA Diamond:92.7% — top score (GPT-4.1: 90.3%)
LiveCodeBench:93.8% — top score (GPT-4.1: 89.4%)
ECI: 147.85 — Ranked 1st among all models
Arena ELO: 1,481 — Top 25

Qwen3.8 27B

GPQA Diamond:90.5% — the top score against 70B-class dense models
ECI:141.90 — #1 among 27B-class models
Local execution:Q4_K_M (18.4GB) → possible on an RTX 4090
Performance:Performance rivaling 70B-class models at 27B parameters

Gemma 4 31B

ECI:128.20 — Google’s top model
Context Window:8K (limited)
Performance:Best ECI among 31B models

DeepSeek-V3 (MoE)

ECI: 112.00 (V3.1: 122.80)
Active Parameters:37B (671B total parameters)
Performance:The highest efficiency among MoE models
Performance:Achieves performance rivalling a 70B-class dense model with 37B active parameters

💻 Section 5: Hardware Requirements

VRAM Number of supported models Representative model
6GB (RTX 3060 6GB) 3 Qwen3-32B-Q2_K, Gemma 3 12B-Q2_K, Llama 3.2 3B
8GB (RTX 3070) 5 Qwen3.8 27B-Q2_K, Gemma 4 31B-Q2_K, Llama 3.1 8B
12GB (RTX 3090) 8 Qwen3.8 27B-Q3_K_M, Gemma 3 27B-Q3_K_M, Llama 3.1 70B-Q2_K
16GB (RTX 4090) 12 Qwen3.8 27B-Q4_K_M, Gemma 4 31B-Q4_K_M, DeepSeek-V3-Q4_K_M
24GB (A100 40GB) 20 Qwen3.8 27B-Q8_0, Gemma 4 31B-Q8_0, MiniMax-M3-Q4_K_M
48GB (A100 80GB) 35 Qwen3.8 Max-Q4_K_M, Kimi K3-Q4_K_M, MiMo-V2.5 Pro-Q4_K_M
80GB+ (H100) 45+ All models
Recommended:When run locallyQ4_K_MQuantization is the most balanced choice. A 27B model runs on 16GB VRAM with almost no performance loss.

📚 Section 6: Data Sources and References

Disclaimer:All benchmark scores were collected from public sources. Real-world model performance may vary by use case.

This pageLocal LLM Hub — What Runs on My Computeris part of. The per-hardware supported-model table is in the hub.