Download it from the ChatGPT screen? That’s not local. Local is getting the weight files onto your own disk and running inference directly on your own GPU. As of 2026, DeepSeek V4, Qwen3.5, and even OpenAI’s gpt-oss all have open weights with permissive Apache/MIT licenses. This post covers installation, download, actually running it, and the realistic line for each of our machines — in one piece.
🌐 · English · 中文 · Español · 한국어 · · English Hub
The 2026 open-weights landscape — what actually runs
| Model (2026) | Structure | Active parameters | Lightweight distribution | Context |
|---|---|---|---|---|
| DeepSeek V4 Flash | MoE 284B | 13B | GGUF 3bit 103GB | 1M |
| DeepSeek V4.1 Flash | MoE 552B | 8B/16B | Ollama is cloud-only | 1M |
| Qwen3.5 9B | Dense 9.7B | All | Ollama Q4_K_M 6.6GB | 256K |
| Qwen3.5 35B-A3B | MoE 35B | 3B | Ollama 24GB (MLX 22GB) | 256K |
| gpt-oss 120B | MoE 117B | 5.1B | Ollama 65GB | 128K |
| gpt-oss 20B | MoE 21B | 3.6B | Ollama 14GB | 128K |
The core is MoE (mixture-of-experts). Even with a large total parameter count, the active part that wakes up per token is only around 3–16B, so a 35B-class Qwen3.5-A3B effectively runs at a 3B price. The distribution sizes in the table are sizes I verified directly today on the ollama.com live tag pages and the Hugging Face unsloth GGUF repos.
Step 1 — Choose a runtime (Ollama / llama.cpp / LM Studio / vLLM)
- Ollama— one-click install on macOS/Windows, model pool, automatic tiered offload, and an OpenAI-compatible API built in. If you’re new, absolutely start here.
- LM Studio— an app for choosing GGUFs and chatting through a graphical interface. For people who dislike terminals.
- llama.cpp— top-tier control (quantization, GPU layers, server flags). Big MoE models like DeepSeek V4 effectively only fit through this door.
- vLLM— for high-throughput servers. Its true value shows in 2-or-more-card llama builds, not a personal desktop.
Step 2 — Installation and first run (verified real-world log)
This is for macOS. On Windows the ollama.com installer .exe follows the same structure.
brew install ollama # or download from ollama.com
ollama pull qwen3.5:9b # downloads 6.6GB of weights
ollama run qwen3.5:9b # start chatting right away
This is the actual log I just ran on my work machine (M5 Max 128GB). Download was 6.6GB; model info after completion:
$ ollama show qwen3.5:9b
architecture qwen35
parameters 9.7B
context length 262144
quantization Q4_K_M
Capabilities: completion, vision, tools, thinking
License: Apache 2.0
Model files stack up in ~/.ollama/models layer by layer, like container images. On my own machine there are currently 14 models totaling 222GB, andollama listanytime to see the ledger.
Step 3 — DeepSeek V4 on my PC — a reality check
As of the official release, V4 Flash (284B, 13B active) ships MIT-licensed weights. But checking ollama.com live today showedthe deepseek-v4.1-flash tag is cloud-only— when you run it, the prompt goes to the Ollama server, not this computer. For truly local, the only path is pulling Hugging Face GGUFs directly with llama.cpp:
# unsloth GGUF, 3bit(103GB) — for 128GB RAM-class machines
llama-cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-IQ3_S \
--temp 1.0 --top-p 1.0 --min-p 0.01
Reality on machines with 64GB or less: V4 Flash’s full 8-bit build is 162GB — a luxury — and even 3-bit (103GB) is borderline. At 32GB or less, give up on the V4 line locally and ask the same question of Qwen3.5 27b — the gap has narrowed a fair bit lately.
Step 4 — What runs on your GPU (directly tied to prices)
| Your rig | Lineup that runs | Dropped lineup |
|---|---|---|
| VRAM 8GB (5050/5060 class) | Qwen3.5 4B·2B, partial offload for gpt-oss 20B | Everything 9B and above |
| VRAM 12GB (5070) | Qwen3.5 9B Q4(6.6GB)+KV headroom, 4B MXFP4 | 13B+ resident |
| VRAM 16GB (5060Ti16/5080) | Qwen3.5 9B + 128K context, gpt-oss 20b | 35B-class resident |
| VRAM 24GB (3090/4090) | Qwen3.5 27b Q4 (17GB), 35b-a3b fits 24GB exactly | 120B class |
| Mac unified memory 128GB (M5 Max) | Qwen3.5 122b-a10b (81GB), gpt-oss 120b (65GB), V4-Flash IQ3 borderline | V4 Pro 1.6T |
To keep the load math simpleVRAM reference table by use caseand just now watchedRTX 5070 12GB limit experimentstogether. The new lineup’s market prices arePrice tracking dashboardis updated daily.
Step 5 — measuring your own machine’s speed directly
curl http://localhost:11434/api/generate -d '{"model":"qwen3.5:9b",
"prompt":"what is mixture-of-experts?","stream":false,
"options":{"temperature":0.2,"num_predict":60}}' | jq '.eval_count, .eval_duration'
# tok/s = eval_count ÷ (eval_duration ÷ 1e9)
Run the same command per card and quantization to build your own t/s table. Other people’s benchmarks won’t transfer directly — RAM clocks and quantization specs differ.
Frequently asked questions
If a download breaks midway, do I have to restart from the beginning?
No. ollama pull resumes per layer, so already-downloaded layers are skipped. Just rerun the same command.
Is it okay to delete a model once downloaded?
ollama rm qwen3.5:9b is deleted and the disk frees up right away. The model list and total actual sizes areollama listanytime to audit — my own machine has 14 models stacked to 222GB, too.
Quantization labels (Q4_K_M, IQ3_S, MXFP4) — which one to pick?
The first button is capacity: do the math and weight GB ≒ parameters × bits ÷ 8. Q4_K_M is a stock phrase — ‘9B means 6.6GB’ is the settled standard — and IQ3_x is a llama.cpp-only format that runs around 3 bits while OS-preserving. When VRAM has room, step up one notch from there.