Benchmarks de LLM locales 2026: tabla comparativa completa de 47 modelos
Fuente:HuggingFace Leaderboard, LMArena Arena ELO, ECI de Benchmarklist.com e informes técnicos oficiales de los modelos. Todos los valores son puntuaciones de benchmarks públicos.
Actualizado:Se actualizará el día 1 de cada mes.Deep Learning Diarioblog.
🌐 · English · 中文 · Español · 한국어 · · Español Hub
🏆 Sección 1: comparación integral de los mejores modelos (ECI + Arena)
Los 10 mejores modelos a septiembre de 2026, según las puntuaciones ECI del benchmark y Arena ELO.
| Puesto | Nombre del modelo | Parámetros | Arquitectura | Vendor | ECI | Arena ELO | GPQA | SWE-bench | HuggingFace |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 Max | 397B | Dense | Qwen | 147.85 | 1,481 | 92.7% | 67.7% | Qwen/Qwen3.8-2.4T-A95B |
| 2 | Qwen3.8 27B | 27B | Dense | Qwen | 141.90 | 1,200 | 90.5% | 58.0% | Qwen/Qwen3.8-27B |
| 3 | Gemma 4 31B | 31B | Dense | 128.20 | — | — | — | google/gemma-4-31b-it | |
| 4 | Mistral Large 3 128B | 128B | Dense | Mistral | 123.00 | — | — | — | mistralai/Mistral-Large-Instruct-2506 |
| 5 | DeepSeek-V3.1 | 671B (37B active) | MoE | DeepSeek | 122.80 | — | — | — | deepseek-ai/DeepSeek-V3.1 |
| 6 | MiniMax-M3 | 428B | Dense | MiniMax | 120.00 | — | — | — | minimax/minimax-m3 |
| 7 | Kimi K3 | 1.1T | MoE | Moonshot | 120.00 | — | — | — | moonshotai/Kimi-K3 |
| 8 | Qwen3.8 Flash-Next | 125B (6B active) | MoE | Qwen | 120.00 | 1,100 | 86.0% | — | Qwen/Qwen3.8-Flash-Next |
| 9 | Command-A+ | 218B | Sparse | Cohere | 118.00 | — | — | — | cohere-ai/command-a-plus |
| 10 | Mixtral 8x22B | 141B (39B active) | MoE | Mistral | 95.00 | — | — | — | mistralai/Mixtral-8x22B-Instruct-v0.1 |
Arena ELO:Ranking de competitividad de modelos calculado con votos de usuarios de LMArena (LMSYS).
GPQA:Benchmark GPQA Diamond: capacidad de resolver problemas de ciencia a nivel de posgrado.SWE-bench:Benchmark de ingeniería de software — capacidad real de arreglar código.
📊 Sección 2: tabla comparativa completa de los 47 modelos
Benchmarks, arquitectura e información pública de todos los LLM locales, clasificados por proveedor.
Serie Qwen (equipo Qwen)
| Nombre del modelo | Parámetros | Context | Release | ECI | Arena Rk | GPQA | SWE-bench | D/L | HuggingFace |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.8 Max | 397B | 262K | Aug 2026 | 147.85 | 21 | 92.7% | 67.7% | 50K | Qwen/Qwen3.8-2.4T-A95B |
| Qwen3.8 27B | 27B | 262K | Aug 2026 | 141.90 | 88 | 90.5% | 58.0% | 6.7M | Qwen/Qwen3.8-27B |
| Qwen3.8 Flash-Next | 125B (6B active) | 256K | Aug 2026 | 120.00 | 149 | 86.0% | — | 1.2M | Qwen/Qwen3.8-Flash-Next |
| Qwen3.6 35B-A3B | 35B (3B active) | 262K | Jul 2026 | 125.70 | 109 | 85.0% | — | 3.3M | Qwen/Qwen3.6-35B-A3B |
| Qwen3.6 27B | 27B | 262K | Jul 2026 | 118.00 | 61 | 89.0% | 50.0% | 2.7M | Qwen/Qwen3.6-27B |
| Qwen3.5 397B-A17B | 397B (17B active) | 262K | Jun 2026 | 108.00 | 81 | 88.0% | — | 305K | Qwen/Qwen3.5-397B-A17B |
| Qwen3-32B | 32B | 128K | Mar 2025 | 110.00 | 208 | — | — | 4.0M | Qwen/Qwen3-32B |
| Qwen3-235B-A22B | 235B (22B active) | 256K | Mar 2025 | 97.00 | 177 | — | — | 371K | Qwen/Qwen3-235B-A22B |
| Qwen3-VL 235B-A22B | 235B (22B active) | 256K | May 2026 | 96.00 | 153 | — | — | 326K | Qwen/Qwen3-VL-235B-A22B-Instruct |
| Qwen2.5-72B Instruct | 72B | 128K | Nov 2025 | 82.00 | 179 | 83.0% | — | 294K | Qwen/Qwen2.5-72B-Instruct |
| Qwen2.5-Coder 32B | 32B | 128K | Nov 2025 | 105.00 | 166 | 84.0% | — | 1.0M | Qwen/Qwen2.5-Coder-32B-Instruct |
| QwQ-32B | 32B | 32K | Jan 2025 | 108.00 | 108 | — | — | 70.6K | Qwen/QwQ-32B |
Serie Meta Llama
| Nombre del modelo | Parámetros | Context | Release | ECI | Arena Rk | GPQA | SWE-bench | D/L | HuggingFace |
|---|---|---|---|---|---|---|---|---|---|
| Llama 4 Maverick 17B-128E | 17B (128 experts) | 128K | Sep 2025 | 110.00 | — | — | — | 8.2K | meta-llama/Llama-4-Maverick-17B-128E-Instruct |
| Llama 4 Scout 17B-16E | 17B (16 experts) | 128K | Sep 2025 | 115.00 | — | — | — | 26.2K | meta-llama/Llama-4-Scout-17B-16E-Instruct |
| Llama 3.3 70B | 70B | 128K | Sep 2025 | 112.00 | — | — | — | — | meta-llama/Llama-3.3-70B-Instruct |
| Llama 3.1 70B | 70B | 128K | Jul 2025 | 108.00 | — | — | — | — | meta-llama/Llama-3.1-70B-Instruct |
| Llama 3.1 8B | 8B | 128K | Jul 2025 | 92.00 | — | — | — | — | meta-llama/Llama-3.1-8B-Instruct |
| Llama 3.2 3B | 3B | 128K | May 2024 | 75.00 | — | — | — | — | meta-llama/Llama-3.2-3B-Instruct |
Serie Google Gemma
| Nombre del modelo | Parámetros | Context | Release | ECI | Arena Rk | GPQA | SWE-bench | D/L | HuggingFace |
|---|---|---|---|---|---|---|---|---|---|
| Gemma 4 31B | 31B | 8K | Jun 2025 | 128.20 | — | — | — | — | google/gemma-4-31b-it |
| Gemma 4 26B-A4B | 26B (4B active) | 8K | Jun 2025 | 118.00 | — | — | — | — | google/gemma-4-26b-A4B |
| Gemma 3 27B | 27B | 128K | Apr 2025 | 113.00 | — | — | — | — | google/gemma-3-27b-it |
| Gemma 3 12B | 12B | 128K | Apr 2025 | 105.00 | — | — | — | — | google/gemma-3-12b-it |
Serie Mistral
| Nombre del modelo | Parámetros | Context | Release | ECI | Arena Rk | GPQA | SWE-bench | D/L | HuggingFace |
|---|---|---|---|---|---|---|---|---|---|
| Mistral Large 3 128B | 128B | 128K | Jul 2025 | 123.00 | — | — | — | — | mistralai/Mistral-Large-Instruct-2506 |
| Mistral Small 3.2 24B | 24B | 128K | Jul 2025 | 115.00 | — | — | — | — | mistralai/Mistral-Small-3.2-24B-Instruct-2506 |
| Mixtral 8x22B Instruct | 141B (39B active) | 32K | Mar 2024 | 95.00 | — | — | — | — | mistralai/Mixtral-8x22B-Instruct-v0.1 |
| Mixtral 8x7B Instruct | 47B (12B active) | 32K | Dec 2023 | 85.00 | — | — | — | — | mistralai/Mixtral-8x7B-Instruct-v0.1 |
Serie DeepSeek
| Nombre del modelo | Parámetros | Context | Release | ECI | Arena Rk | GPQA | SWE-bench | D/L | HuggingFace |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-V3.1 | 671B (37B active) | 128K | Jan 2025 | 122.80 | — | — | — | — | deepseek-ai/DeepSeek-V3.1 |
| DeepSeek-V3.2 | 685B (37B active) | 128K | Jan 2025 | 115.00 | — | — | — | — | deepseek-ai/DeepSeek-V3.2 |
| DeepSeek-R1 | 671B (37B active) | 128K | Dec 2024 | 113.00 | — | — | — | — | deepseek-ai/DeepSeek-R1 |
| DeepSeek-V3 | 671B (37B active) | 128K | Dec 2024 | 112.00 | — | — | — | — | deepseek-ai/DeepSeek-V3 |
| DeepSeek-Coder-V2 | 16B (27B active) | 128K | May 2024 | 105.00 | — | — | — | — | deepseek-ai/DeepSeek-Coder-V2 |
Otros modelos principales
| Nombre del modelo | Parámetros | Context | Release | ECI | Arena Rk | GPQA | SWE-bench | D/L | HuggingFace |
|---|---|---|---|---|---|---|---|---|---|
| GLM-4.7 (Zhipu AI) | 355B | 128K | Jul 2025 | 115.00 | — | — | — | — | THUDM/glm-4-9b-chat |
| GLM-4 Plus (Zhipu AI) | 176B | 128K | May 2025 | 110.00 | — | — | — | — | THUDM/glm-4-9b-chat |
| Command-A+ (Cohere) | 218B | 128K | Aug 2025 | 118.00 | — | — | — | — | cohere-ai/command-a-plus |
| Command-R+ (Cohere) | 104B | 128K | Sep 2024 | 108.00 | — | — | — | — | cohere-ai/command-r-plus |
| Phi-4 (Microsoft) | 14B | 16K | Mar 2025 | 98.00 | — | — | — | — | microsoft/phi-4 |
| Phi-3 Medium (Microsoft) | 14B | 4K | Mar 2024 | 82.00 | — | — | — | — | microsoft/phi-3-medium-4k-instruct |
| Nemotron-4 340B (NVIDIA) | 340B | 128K | Jun 2025 | 108.00 | — | — | — | — | nvidia/Nemotron-4-340B-Instruct |
| Nemotron-3 Super 120B (NVIDIA) | 120B (12B active) | 128K | Jun 2025 | 105.00 | — | — | — | — | nvidia/Nemotron-3-Super-120B-A12B |
| OLMo-3 32B Think (Allen AI) | 32B | 128K | Jul 2025 | 96.00 | — | — | — | — | allenai/OLMo-3-32B-Think |
| OLMo-3 32B (Allen AI) | 32B | 128K | Jul 2025 | 94.00 | — | — | — | — | allenai/OLMo-3-32B |
| Step-3 (StepFun) | 198B | 128K | Aug 2025 | 108.00 | — | — | — | — | stepfun-ai/step-3 |
| Kimi K3 (Moonshot) | 1.1T | 128K | Jul 2025 | 120.00 | — | — | — | — | moonshotai/Kimi-K3 |
| MiMo-V2.5 Pro (Xiaomi) | 1.02T | 128K | Aug 2025 | 115.00 | — | — | — | — | xiaomi-ai/mimo-v2.5-pro |
| InternLM 2.5 20B (InternLM) | 20B | 128K | Aug 2024 | 92.00 | — | — | — | — | internlm/internlm2.5-20b |
| Solar 10.7B (Upstage) | 10.7B | 128K | Sep 2024 | 90.00 | — | — | — | — | upstage/SOLAR-10.7B-Instruct-v1.0 |
⚡ Sección 3: Tabla comparativa de cuantización
Formatos de cuantización disponibles y requisitos de memoria por modelo.
Variantes de cuantización de Qwen3.8 27B
| Formato | Número de parámetros | Memoria | Descarga | URL |
|---|---|---|---|---|
| safetensors (Original) | 27.4B | 55.8 GB | 2.82M | Qwen/Qwen3.8-27B |
| Q4_K_M | 27.4B | 18.4 GB | — | Qwen/Qwen3.8-27B-Q4_K_M |
| Q4_K_S | 27.4B | 15.7 GB | — | Qwen/Qwen3.8-27B-Q4_K_S |
| Q8_0 | 27.4B | 30.5 GB | — | Qwen/Qwen3.8-27B-Q8_0 |
| Q5_1 | 27.4B | 22.4 GB | — | Qwen/Qwen3.8-27B-Q5_1 |
| Q5_K_M | 27.4B | 21.1 GB | — | Qwen/Qwen3.8-27B-Q5_K_M |
| Q3_K_L | 27.4B | 14.5 GB | — | Qwen/Qwen3.8-27B-Q3_K_L |
| Q3_K_M | 27.4B | 13.2 GB | — | Qwen/Qwen3.8-27B-Q3_K_M |
| Q2_K | 27.4B | 10.5 GB | — | Qwen/Qwen3.8-27B-Q2_K |
🎯 Sección 4: resumen de logros clave
Qwen3.8 Max
GPQA Diamond:92,7% — máxima puntuación (GPT-4o: 88,9%, Claude Opus 4.6: 85,0%)
SWE-bench Verified:67.7% — mejor puntuación (GPT-4.1: 57.2%, Claude Opus 4.6: 54.5%)
LiveCodeBench:93.8% — mejor puntuación (GPT-4.1: 89.4%)
SWE-bench Full:63.0% — mejor puntuación
MMLU-Pro:87.8% — la mejor puntuación (GPT-4.1: 85.6%)
HumanEval:91,8% — máxima puntuación (GPT-4.1: 90,4%)
GPQA Diamond:92,7% — máxima puntuación (GPT-4.1: 90,3%)
LiveCodeBench:93.8% — mejor puntuación (GPT-4.1: 89.4%)
ECI: 147.85 — N.º 1 entre todos los modelos
Arena ELO: 1,481 — Top 25
Qwen3.8 27B
GPQA Diamond:90.5% — la mejor puntuación frente a densos de categoría 70B
ECI:141,90 — 1.º entre los modelos de clase 27B
Ejecución local:Q4_K_M (18.4GB) → posible en una RTX 4090
Rendimiento:Rendimiento a la altura de modelos de 70B con 27B de parámetros
Gemma 4 31B
ECI:128.20 — el modelo estrella de Google
Context Window:8K (limitado)
Rendimiento:Mejor ECI entre los modelos 31B
DeepSeek-V3 (MoE)
ECI: 112.00 (V3.1: 122.80)
Active Parameters:37B (671B de parámetros totales)
Rendimiento:La mayor eficiencia entre los modelos MoE
Rendimiento:Rendimiento comparable a un modelo denso de clase 70B con 37B de parámetros activos
💻 Sección 5: requisitos de hardware
| VRAM | Modelos compatibles | Modelo representativo |
|---|---|---|
| 6GB (RTX 3060 6GB) | 3 | Qwen3-32B-Q2_K, Gemma 3 12B-Q2_K, Llama 3.2 3B |
| 8GB (RTX 3070) | 5 | Qwen3.8 27B-Q2_K, Gemma 4 31B-Q2_K, Llama 3.1 8B |
| 12GB (RTX 3090) | 8 | Qwen3.8 27B-Q3_K_M, Gemma 3 27B-Q3_K_M, Llama 3.1 70B-Q2_K |
| 16GB (RTX 4090) | 12 | Qwen3.8 27B-Q4_K_M, Gemma 4 31B-Q4_K_M, DeepSeek-V3-Q4_K_M |
| 24GB (A100 40GB) | 20 | Qwen3.8 27B-Q8_0, Gemma 4 31B-Q8_0, MiniMax-M3-Q4_K_M |
| 48GB (A100 80GB) | 35 | Qwen3.8 Max-Q4_K_M, Kimi K3-Q4_K_M, MiMo-V2.5 Pro-Q4_K_M |
| 80GB+ (H100) | 45+ | Todos los modelos |
📚 Sección 6: fuentes de datos y referencias
- HuggingFace Leaderboard: open-llm-leaderboard
- LMArena (LMSYS): lmarena.ai
- Benchmarklist.com: benchmarklist.com — ECI scores
- Qwen3.8 Technical Report: qwen.ai/blog
- Modelscope: modelscope.cn
- GPQA Diamond: github.com/idavidrein/gpqa
- SWE-bench: github.com/SWE-bench/SWE-bench
Esta páginaHub de LLM locales — qué corre en mi ordenadores parte de. La tabla de modelos compatibles por hardware está en el hub.