GPU 시세 · VRAM · 온디바이스 AI

Download it from the ChatGPT screen? That’s not local. Local is getting the weight files onto your own disk and running inference directly on your own GPU. As of 2026, DeepSeek V4, Qwen3.5, and even OpenAI’s gpt-oss all have open weights with permissive Apache/MIT licenses. This post covers installation, download, actually running it, and the realistic line for each of our machines — in one piece.

🌐 · English · 中文 · Español · 한국어 · · English Hub

The 2026 open-weights landscape — what actually runs

Model (2026) Structure Active parameters Lightweight distribution Context
DeepSeek V4 Flash MoE 284B 13B GGUF 3bit 103GB 1M
DeepSeek V4.1 Flash MoE 552B 8B/16B Ollama is cloud-only 1M
Qwen3.5 9B Dense 9.7B All Ollama Q4_K_M 6.6GB 256K
Qwen3.5 35B-A3B MoE 35B 3B Ollama 24GB (MLX 22GB) 256K
gpt-oss 120B MoE 117B 5.1B Ollama 65GB 128K
gpt-oss 20B MoE 21B 3.6B Ollama 14GB 128K

The core is MoE (mixture-of-experts). Even with a large total parameter count, the active part that wakes up per token is only around 3–16B, so a 35B-class Qwen3.5-A3B effectively runs at a 3B price. The distribution sizes in the table are sizes I verified directly today on the ollama.com live tag pages and the Hugging Face unsloth GGUF repos.

Step 1 — Choose a runtime (Ollama / llama.cpp / LM Studio / vLLM)

  • Ollama— one-click install on macOS/Windows, model pool, automatic tiered offload, and an OpenAI-compatible API built in. If you’re new, absolutely start here.
  • LM Studio— an app for choosing GGUFs and chatting through a graphical interface. For people who dislike terminals.
  • llama.cpp— top-tier control (quantization, GPU layers, server flags). Big MoE models like DeepSeek V4 effectively only fit through this door.
  • vLLM— for high-throughput servers. Its true value shows in 2-or-more-card llama builds, not a personal desktop.

Step 2 — Installation and first run (verified real-world log)

This is for macOS. On Windows the ollama.com installer .exe follows the same structure.

brew install ollama        # or download from ollama.com
ollama pull qwen3.5:9b     # downloads 6.6GB of weights
ollama run qwen3.5:9b      # start chatting right away

This is the actual log I just ran on my work machine (M5 Max 128GB). Download was 6.6GB; model info after completion:

$ ollama show qwen3.5:9b
  architecture        qwen35
  parameters          9.7B
  context length      262144
  quantization        Q4_K_M
  Capabilities: completion, vision, tools, thinking
  License: Apache 2.0

Model files stack up in ~/.ollama/models layer by layer, like container images. On my own machine there are currently 14 models totaling 222GB, andollama listanytime to see the ledger.

Step 3 — DeepSeek V4 on my PC — a reality check

As of the official release, V4 Flash (284B, 13B active) ships MIT-licensed weights. But checking ollama.com live today showedthe deepseek-v4.1-flash tag is cloud-only— when you run it, the prompt goes to the Ollama server, not this computer. For truly local, the only path is pulling Hugging Face GGUFs directly with llama.cpp:

# unsloth GGUF, 3bit(103GB) — for 128GB RAM-class machines
llama-cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-IQ3_S \
  --temp 1.0 --top-p 1.0 --min-p 0.01

Reality on machines with 64GB or less: V4 Flash’s full 8-bit build is 162GB — a luxury — and even 3-bit (103GB) is borderline. At 32GB or less, give up on the V4 line locally and ask the same question of Qwen3.5 27b — the gap has narrowed a fair bit lately.

Step 4 — What runs on your GPU (directly tied to prices)

Your rig Lineup that runs Dropped lineup
VRAM 8GB (5050/5060 class) Qwen3.5 4B·2B, partial offload for gpt-oss 20B Everything 9B and above
VRAM 12GB (5070) Qwen3.5 9B Q4(6.6GB)+KV headroom, 4B MXFP4 13B+ resident
VRAM 16GB (5060Ti16/5080) Qwen3.5 9B + 128K context, gpt-oss 20b 35B-class resident
VRAM 24GB (3090/4090) Qwen3.5 27b Q4 (17GB), 35b-a3b fits 24GB exactly 120B class
Mac unified memory 128GB (M5 Max) Qwen3.5 122b-a10b (81GB), gpt-oss 120b (65GB), V4-Flash IQ3 borderline V4 Pro 1.6T

To keep the load math simpleVRAM reference table by use caseand just now watchedRTX 5070 12GB limit experimentstogether. The new lineup’s market prices arePrice tracking dashboardis updated daily.

Step 5 — measuring your own machine’s speed directly

curl http://localhost:11434/api/generate -d '{"model":"qwen3.5:9b",
  "prompt":"what is mixture-of-experts?","stream":false,
  "options":{"temperature":0.2,"num_predict":60}}' | jq '.eval_count, .eval_duration'
# tok/s = eval_count ÷ (eval_duration ÷ 1e9)

Run the same command per card and quantization to build your own t/s table. Other people’s benchmarks won’t transfer directly — RAM clocks and quantization specs differ.

Frequently asked questions

If a download breaks midway, do I have to restart from the beginning?

No. ollama pull resumes per layer, so already-downloaded layers are skipped. Just rerun the same command.

Is it okay to delete a model once downloaded?

ollama rm qwen3.5:9b is deleted and the disk frees up right away. The model list and total actual sizes areollama listanytime to audit — my own machine has 14 models stacked to 222GB, too.

Quantization labels (Q4_K_M, IQ3_S, MXFP4) — which one to pick?

The first button is capacity: do the math and weight GB ≒ parameters × bits ÷ 8. Q4_K_M is a stock phrase — ‘9B means 6.6GB’ is the settled standard — and IQ3_x is a llama.cpp-only format that runs around 3 bits while OS-preserving. When VRAM has room, step up one notch from there.