GPU 시세 · VRAM · 온디바이스 AI

The limit for running LLMs on a laptop is set by VRAM, not the CPU. Simple arithmetic: take usable VRAM as nameplate × 0.88, subtract Q4 weights (~0.61GB per billion parameters) and the KV cache, and you have your answer — this post tables the model sizes that actually run in the 8GB, 12GB, 16GB and 24GB bands. Add September 2026 Danawa laptop prices and “how many B fit my budget” fits in one paragraph.

🌐 · English · 中文 · Español · 한국어 · · English Hub

The first question for laptop LLMs isn’t the CPU, it’s VRAM

In quantized local LLM inference, the bottlenecks in order are: load capacity, memory bandwidth, cores. If the model doesn’t fit in VRAM, before any speed question it either won’t run at all or spills into system RAM at 5~10x slower speed. And laptop GPUs with the same name often have different VRAM. The 2026 RTX 5070 laptop ships with 8GB (128-bit, ~384GB/s) by default, but memory supply issues led to a 12GB version launching in April 2026, and both models are now sold side by side. Danawa search results mix “VRAM: 8GB” and “VRAM: 12GB” labels, so even with an identical model name, always check the spec sheet.

Three numbers in the VRAM formula: 0.88, 0.61, KV cache

The calculation has three steps. First, on Windows only about 88% of nominal VRAM is actually usable due to CUDA overhead. Second, for llama.cpp-style Q4_K_M quantization, weights are roughly 0.61GB per 1B parameters. Third, add the KV cache to hold the context: bytes per token = 2(K,V) x layers x KV heads x head size (128) x 2 bytes (FP16), multiplied by the context length. For example, Llama 3.1 8B (32 layers, 8 KV heads) works out to 2 x 32 x 8 x 128 x 2 = 131,072 bytes/token, about 1.0GB for an 8K context. Add a 1GB buffer for the OS and runtime and you get the final requirement.

Band Nominal VRAM Usable capacity (x0.88) Comfortable for daily use at 8K
Entry-level 8 GB 7.0 GB 8B Q4 is the limit, no context headroom
Flagship pick 12 GB 10.6 GB Up to 12B Q4; 14B gets truncated
Near-flagship 16 GB 14.1 GB 14B Q4 + 8K context possible
Top tier 24 GB 21.1 GB Up to 24B·30B MoE

Laptop GPU VRAM map: same name, different capacity

The official VRAM for the 2026 RTX 50-series laptop lineup is 5070 8GB (128-bit), 5070 Ti 12GB (192-bit), 5080 16GB (256-bit), 5090 24GB (256-bit). A 5070 12GB variant has been added. Framework module pricing is a real example of how big the VRAM premium is: the same 5070 module costs $699 at 8GB and $1,199 at 12GB — an extra $500 for 4GB of VRAM. When buying a laptop, don’t just look at the GPU name; filter the VRAM field directly in Danawa’s spec filter.

Laptop GPU VRAM Bus width Runnable line at 8K
RTX 5070 (variants included) 8 / 12 GB 128-bit 8B confirmed; for 12B only the 12GB build
RTX 5070 Ti 12 GB 192-bit Up to 12B Q4
RTX 5080 16 GB 256-bit 14B Q4 + long context
RTX 5090 24 GB 256-bit 24B, 30B MoE class

Measured VRAM requirements by LLM model size on laptops

The weight GB below are measured file sizes of GGUF files (Q4_K_M class, as of 2026-09-27) I actually ran on my M5 Max MacBook with llama.cpp. The Q4 8K KV is computed from the formula above, and the required total = weights + KV + 1GB buffer.

Model (Q4) Weights KV (8K) Total Minimum VRAM tier
Llama 3.1 8B 4.9 GB 1.0 GB 6.9 GB 8GB (7.0 usable, exact fit)
12B-class GQA 7.3 GB 1.25 GB 9.6 GB 12GB
Phi-4 14B 9.1 GB 1.56 GB 11.7 GB 16GB
gpt-oss 20B (MoE) 13.8 GB 0.75 GB 15.6 GB 24GB
Mistral Small 24B 14.3 GB 1.25 GB 16.6 GB 24GB
Qwen3 Coder 30B (MoE) 18.6 GB 0.75 GB 20.4 GB 24GB (last column)
70B-class 42.5 GB 2.5 GB 46.0 GB Not possible on a single laptop card

Here’s how to read it. An 8B runs even on an 8GB laptop, but with 6.9GB against 7.0GB usable there’s only 0.1GB of headroom, so context has to shrink to 4K or less. 12B is the first slot of the 12GB band; 14B doesn’t fit the 14.1GB band, so it needs 16GB. 70B doesn’t hold in any laptop-card band to begin with.

What if it won’t run: quantization, MoE, RAM offload

The first lever is stronger quantization. Dropping Q4 to Q3 or IQ2 cuts weights by 25~45%, but sentence comprehension visibly degrades. The second is MoE models. gpt-oss 20B has 13.8GB of weights but only 3.6B active parameters, so even with RAM offload slowing token speed, the perceived slowdown is smaller than with a dense model. The third is RAM offload: DDR5-5600 dual-channel bandwidth is about 90GB/s, a quarter of a laptop GPU’s, so pushing weights into RAM drops speed 5~10x. In other words, “it runs” and “it’s usable” are different sentences. Filling RAM to at least 32GB is the minimum to leave room for offloading.

Buying criteria by budget: student · lab · office worker

Students (laptop budget around ₩2M): the RTX 5070 8GB laptop is the entry line. On Danawa, the ASUS TUF A18 5070 8GB build is listed at ₩2,099,000 with the coupon applied, and the MSI Crosshair 16 5070 8GB sits in the ₩2,729,000 range. Realistic uses are 8B Q4 inference and a coding assistant for classes. Lab (₩3M range): for the same money, pick a 5070 Ti 12GB or a 5070 12GB variant. Once 12B Q4 fits, practicality jumps to a whole new tier. If it can double as a desktop, an RTX 5070 12GB card is ₩1,383,800 (Danawa median, as of September 28), so the total including the base unit comes out cheaper. Office workers / high budget: a 5080 16GB or 5090 24GB laptop; if mobility matters less, a used 3090 24GB desktop still wins on cost per GB. The issue that even the same name can differ in performance by up to 40% depending on TGPLaptop GPU TGP fully masteredcovered it.

The fork in the road versus desktop

The laptop 5070 8GB and desktop 5070 12GB share a name but are different machines. The desktop 5070 has a wider 192-bit bus, so even the same 8B inference runs clearly faster in tokens, and it carries 4GB more VRAM. The practical limit of a 12GB cardRTX 5070 12GB VRAM: real-world limitscomputed from, and the minimum/recommended amounts covering all use cases areDeep-learning GPU VRAM reference table by use casesee it together. Unless portability is a must, the structure where desktop wins on price-per-GB hasn’t changed in 2026 either.

Frequently asked questions

Does Llama 3.1 8B definitely run on an 8GB laptop?

It runs. But with 7.0GB usable and 4.9GB weights + 1.0GB KV + buffer, it fits exactly — you’d need to cut context to 4K and close background apps. For comfort, go for the 12GB band.

Will a 70B ever work on a laptop?

It won’t fit on a single card. Q4 weights alone are 42.5GB — tight even split across two 24GB cards. RAM offload can run it, but on laptop DDR5 bandwidth token speed drops to single digits.

How do I choose between the 5070 12GB variant and the 5070 Ti 12GB?

Both are 12GB, but the Ti has a 192-bit bus width, so higher bandwidth gives it an edge in long context. If the budget is the same, go Ti; if the price gap is large, the 12GB variant is reasonable. For both, first check the price difference against the 5080 16GB.

If I upgrade RAM to 64GB, won’t bigger models run?

It “runs” — but with weights in RAM you’re using RAM bandwidth instead of GPU bandwidth, so speed drops 5–10x. Only families with small active parameters, like MoE, stay at usable speeds.

Sources: NVIDIA official spec sheets, Framework public pricing, Danawa actual listings (collected Sep 27–28, 2026), and my own M5 Max llama.cpp-measured GGUF sizes. Prices and specs change, so check the latest information before buying.