GPU 시세 · VRAM · 온디바이스 AI

The 5060 Ti 16GB is the cheapest 16GB card in the 2026 local LLM market. On Windows, OS+driver eat about 1GB, so the practical weight budget is 13–14.5GB (with KV reduced to q4_0 and a shortened context). All sizes in the table below were measured today from live Hugging Face files.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Model Build Weights Speed Source/notes
Qwen3.6-35B-A3B IQ3_XXS 13.2 GB 47~51 tok/s (measured at 160K context) njannasch.dev measured — 15,847MiB used including KV q4_0 962MiB
Qwen3-Coder-30B-A3B Q3_K_M 14.7 GB 30B-class #1 recommendation r/LocalLLaMA 5060 Ti thread — ’30B still wins’
gpt-oss-20b Q4_K_M 11.6 GB 242.8 tok/s on the 5080 (upstream GGUF) techfuelhq measured — the 5060 Ti is lower than this
Qwen3.5-9B Q4_K_M 6.8 GB estimated 40~51 tok/s noted: modelfit.io estimate 51 tok/s for 8B class
Llama-3.1-8B Q8_0 8.5 GB estimated ~50 tok/s modelfit.io — 8B-class 51 tok/s
Phi-4 Q6_K 12.0 GB Est. ~30 tok/s One bit above modelfit.io’s 14B Q4 32 tok/s baseline
DeepSeek-R1-Distill-8B Q4_K_M 4.9 GB Est. 45~55 tok/s Based on 8B dense
Mistral-Small-3.2-24B Q3_K_M 11.5 GB Not measured 14GB line including KV — fits, but just barely
Gemma 4 26B-A4B QAT UD-Q4_K_XL 14.2 GB Unmeasured — the most likely new entrant QAT, so top-tier 4-bit quality — KV q8 essential

What doesn’t fit

  • Qwen3.8-27B Q4_K_M 16.5GB — over capacity on weights alone. With two cards (5060 Ti ×2) there are reports of 130 tok/s measured (r/LocalLLaMA).
  • Muse-Glimmer-30B Q4_K_M 16.8GB, Laguna-XS-2.1 Q4_K_M 19.6GB — 24GB-card territory.
  • Dense 30B or above — only MoE (A3B-class) is realistic.

Runtime choice

Public benchmarks on the same card: llama.cpp 77.0 / LM Studio 76.8 / Ollama 69.1 tok/s (Qwen2.5-Coder-7B Q4, inventivehq). LM Studio for GUI, Ollama for an API service — but Ollama has a context-slider default trap —3-way runtime comparisonNote. One batch/KV setting change outweighs swapping models.

Longer recipe list in the community-accumulated repoclub-5060tiis here. The laptop 5060 Ti (lower TGP) isLaptop GPU TGP guidefirst. The VRAM calculation formula isVRAM guide by use.

Written 2026-09-28 · file sizes measured from live Hugging Face listings · speeds only from sourced public measurements and my own · no local runs ·View the full hub · Quantization ladder, 104 builds · Generation-speed chart, 47 models