GPU 시세 · VRAM · 온디바이스 AI

I read three articles and redid the math on my own calculator to rewrite this. Facts and figures come from the originals and market data; the formulas and conclusions come from this article. Catch any wrong calculations in the comments.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Serverless vs my 4090: the math is a step ahead of the article

1. ‘Elasticity and Cost’ — the two faces of serverless AI

Here’s how the article frames it: load swings → serverless; steady 24/7 → dedicated hardware. Formally correct. But plug in 2026 Korean price tables and the conclusion shifts. On Joonggonara, the RTX 4090’s average market price is ₩3.56M. Memory shortages pushed new-card prices up and used ones followed — a market where the buyer’s advantage is over.

Here’s the math I ran. 3.56 million won depreciated over 36 months is 99,000 won/month. Power: 0.35kW average for inference, running 24 hours, is 252kWh/month; at the top residential tier rate of 250 won, that’s 63,000 won/month. Total: around 160,000 won/month. On the same card, a 13B-class quantized model at 25 tokens per second yields a conservative 65 million tokens/month. About 2 won per million tokens. For simple document classification and code assist, the performance headroom is still there.

Here’s the conclusion the article doesn’t state. If you pay more than 200,000 KRW/month for metered inference APIs, then for any workload a 13B-class model can replace, local is already overwhelmingly cheaper. So the real reason to buy serverless isn’t cost but frontier-model performance and zero ops. Change the question: ask for the ratio, not the price. What percentage of your workload can be pushed to a smaller model? Once that ratio exceeds 70, the bill is leaking money, not strategy; teams under 30 are right to stay serverless. The inflection point isn’t model size but the difficulty distribution of the work.

One variable the article omits. In a cycle where HBM and DDR rise together, the GPU stopped being a consumable and became an asset with residual value. A hedge line has been added to the depreciation table of ownership. I’ll also state the falsification condition: if in a year the used 4090 price falls below ₩3.56M, this logic is wrong.

Nvidia · MediaTek: not a move to kill a rival, but to grow the board

2. Nvidia secures AI-infrastructure leadership with $3.5B investment in MediaTek

Nvidia buys $3.5 billion in MediaTek convertible bonds; MediaTek adopts NVLink Fusion. Three-track cooperation: rack-scale infrastructure, PC SoCs, and automotive. Commentators talk lock-in; I say the stake math comes first.

Instead of blocking hyperscalers’ custom-silicon escape from GPUs, Nvidia wove them into its own interconnect. Customers keep buying NVLink and rack networking even while building their own chips, and the chip MediaTek will design still takes HBM. It’s a deal that builds a bridge AMD can’t cross in logic competition while growing the whole pie of memory demand. From SK hynix’s perspective the memory pie grows, so there’s nothing to lose — indeed today’s closing price rose more than Samsung’s.

The evidence shows up in next quarter’s supply-chain report. If MediaTek’s custom-accelerator orders are confirmed, TSMC utilization will rise, but HBM device counts don’t care which platform. My take: this deal is memory-volume news, not logic news.

Apple Watch always-on listening: not a battery constant — an NPU constant problem

3. Always listening, never recording — Apple’s AI privacy details revealed

The point is two sentences. The S11 processor handles voice always-on without using the battery. It processes only text, leaving no raw audio behind. But the real bottleneck of this design isn’t the battery.

Without keeping the raw audio, a live rewind that replays what you said 15 seconds ago is fundamentally impossible. ‘Keeping only text’ means cutting upstream; what holds it up is an ultra-low-power always-on audio pipeline — ultimately NPU and SRAM bandwidth design. The moment Apple makes this work on the wrist, the fight left for rivals isn’t speech recognition accuracy but inference per watt.

The fallout for developers is even bigger. The moment always-on inference runs on your wrist, the voice assistant’s standby indicator stops being a cost and becomes table stakes, and voice-service providers built on cloud round-trips have to redo their cost math. For those of us burning local LLMs at our desks, what matters is this being the first commercial case where power constraints beat software constraints.