GPU 시세 · VRAM · 온디바이스 AI

Laptop GPU TGP is the power ceiling that can split performance by up to 40 percentage points even on the same chip. Buy on the GPU name alone and you can end up with a thin ultrabook training at half the speed of a high-end gaming laptop. Below is a table covering how to check TGP, Dynamic Boost and thermals, and what to prioritize for ML training.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Why TGP matters so much

TGP (Total Graphics Power) is the power ceiling (W) the laptop maker allows the GPU. The same chip may run at 35W, or at 115W with dynamic boost pulling up to 140W. Clocks and voltages are decided within this ceiling, so TGP determines real-world performance before the GPU model name does.

Why two identical RTX 5070 laptops differ by 40% in performance

Compare the top and bottom of the table below: perceived performance vs desktop splits from 40–55% to 80–95% — a 40-point gap (95-55=40). Even GPUs with the same name run different voltage/clock operating ranges when their power limits differ, and that difference shows up directly in frames and training speed.

TGP band Typical laptop Perceived vs desktop RTX 7B INT4 inference throughput (relative) Example class
35~60W Thin, light ultrabook (MAX-Q-class low power) 40~55% About 15~25 tok/s LG Gram-class ultralight
80~100W Budget gaming 60~75% About 30–45 tok/s Legion 5 / ROG Strix G class
115~140W (with Dynamic Boost) High-end gaming/workstation 80~95% about 55–80 tok/s Legion Pro/MSI Titan class

Throughput bands are relative figures based on public benchmark averages, assuming a 7B INT4 quantized model at batch-1 inference; actual results vary by model and cooling.

Where to check GPU power (W)

The rule is to first find the ‘GPU Maximum Power’ item on the manufacturer’s official spec sheet. If the number is vague, check the actual sustained power in a review’s GPU-Z load screenshot, and cross-reference with a benchmark like Notebookcheck that compares the same GPU across different TGP models. Also always check whether the listed figure includes Dynamic Boost.

Dynamic Boost and heat: why benchmarks collapse in long training runs

Dynamic Boost shifts unused CPU power headroom to the GPU when CPU load is low — the 115W+25W=140W rows in the table above are built that way. The problem is duration. Benchmarks last seconds to minutes, so the boost holds, but ML training runs for hours, saturating the chassis with heat until thermal throttling kicks in, and the sustained clocks then settle 10~20% below peak. That’s why a high-TGP laptop picked on short benchmark scores disappoints in long training runs.

What to do with a low-TGP laptop

If your setup is ‘training on server/Colab, laptop for development + lightweight inference,’ a low-TGP ultrabook is actually the reasonable choice. It wins on battery, noise, and weight, and with 32GB+ RAM combined with quantized inference it’s perfectly usable as a dev environment. Just note the VRAM capacity ceiling is a separate issue that sets the upper bound for local LLMs, soRTX 5070 12GB VRAM: real-world limitsis worth reading alongside.

In the M5 MacBook era, is an eGPU the answer

To cut to the chase: an eGPU for CUDA isn’t a realistic option. Apple Silicon MacBooks can’t use external GPUs for training, and even a desktop-class eGPU delivers poor efficiency relative to enclosure and cable costs — for the same budget, you’re better off raising the laptop’s own TGP and cooling. The safe play is to keep the MacBook in its lane: Apple GPU inference only.

Laptop buying checklist for ML training

  1. Check GPU power (W) on the manufacturer’s official spec sheet — distinguish whether the figure includes Dynamic Boost
  2. Performance compared against other models with the same GPU on review benches (see Notebookcheck)
  3. Cooling check under sustained load — look at clocks after 30 minutes, not peak
  4. Check 32GB RAM and a free SSD slot — upgradability decides the machine’s life more than the GPU

Frequently asked questions

Does a local LLM work on an 8GB laptop GPU?

7–8B-class INT4 quantized models will run. But as context grows, the KV cache eats VRAM and answer length gets capped, and on low-TGP ultrabooks per-token speed drops as well. If you summarize lots of documents, I’d recommend 12GB or more.

When the spec sheet lists two power figures, which one do you trust?

When both default TGP and maximum power including Dynamic Boost are listed, judge sustained performance by the default TGP. On all-core sustained workloads like training, boost headroom rarely materializes.

Can you fine-tune on a low-TGP laptop?

Lightweight fine-tuning like LoRA works even at the 100W class, but longer training time means heat accumulates. If you plan to train constantly, pick a 115W-plus model — or keep the laptop for development and push training to a server; that’s the better deal per hour.

Source: performance ratios are averages from public benchmarks (Notebookcheck, etc.); actual results vary by model and cooling. The tok/s bands are relative figures assuming 7B INT4 quantization and batch-1 inference; no pricing is included.