GPU 시세 · VRAM · 온디바이스 AI

I read all six myself. Case studies are ads, but numbers hide inside the ads. I pulled out just those numbers and converted them into Korean labor costs and unit economics.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Converting the Codex case study to Korean labor costs

1. Proaction boosts sales 60% and saves 75+ hours with Codex

This is a talk from vehicle-fleet-management startup Proaction. A non-developer co-founder fed sales call recordings, emails, and spreadsheets into Codex and produced a demo that mirrors the customer’s vehicle configuration exactly in 30 to 45 minutes, saved 40–60 engineering hours a month, and grew revenue 60%.

The real number isn’t 60% but time. A Korean developer’s 60M KRW salary converts to roughly 30,000 KRW/hour fully loaded with the four major insurances and overhead. 50 hours a month is 1.5M KRW. AI tool seat fees are a toy next to that. But the shock is elsewhere: when demo production drops from the usual two days to 30 minutes, that gap is quoting speed itself — pipeline volume itself.

What Korean B2B orgs should steal isn’t the tools but the structure. While requirements definition and demo production live in separate departments, this meeting will keep losing quoting speed. If you can’t restructure departments, at least build a board the sales executive uses personally.

Common keywords across the three GPT-6 Astra cases: context is an asset, reasoning steps are cost

2. How Harvey turned legal context into draft quality

Among what legal AI Harvey unveiled, the memory panel catches my eye more than model performance. It’s a mechanism that stores per-lawyer rules — numbered lists, EDGAR-first, priority colors — and enforces them in drafts. That structure matters more than claims of a longer context window.

Because everyone buys models, but preference data stays proprietary. In three years the legal-AI gap won’t be about models — it’ll be about who owns each firm’s accumulated preferences. The same question goes to Korean law firms and legaltech: are the corrections dictated by lawyers today being logged as data, or just fixed and forgotten? Only what’s logged accumulates.

3. invideo: the math behind a 3x jump in color-grading success rate

Video-editing agent invideo announced a 3x jump in color-grading success rates. What sticks in the calculator more than that sentence: the same job finished in far fewer inference steps.

For agent products, inference steps are literally cost of goods sold. Cut a 20-step task to 4 steps and that service’s gross margin jumps by the step-reduction factor. Read this series not as a performance announcement but as a cost-structure overhaul. If you run your own agents, count today’s average step count per task. In six months’ model-generation comparison tables, that number will prove unit economics for you.

4. Two numbers Airbnb shared while announcing expanded model access

Two numbers matter. On strategy-doc work — not coding — revisions that took 20+ tries with other models took 3~4. And shipped features grew 80% year-over-year without adding headcount.

The first number changes the order of enterprise contract negotiations: you should bet on success rate per attempt, not price per seat, because you can’t pay for 25 attempts on a job that takes 4. The line item in the contract should be passes, not users. The second number is the logic of a hiring-frozen organization: when hiring is blocked, output per head is the hire, and 80% is the proof.

ChatGPT ads: Korea is already not in the game

5. ChatGPT Ads expands to Southeast Asia and Taiwan

Adding seven Southeast Asian markets is this round’s news, but South Korea was already in an earlier rollout. Ads appear only on the Free and Go plans; Plus, Pro, and Enterprise are ad-free. On the agency side, dentsu, Publicis, and WPP are on board.

The essence is that purchase context is moving from search-box keywords to conversation paragraphs. Even for the same product inquiry, chat utterances are more than twice as long, and ad creative and landing-page length have to grow with it. Search inventory is saturated so bid prices are already high; conversational is early so it’s relatively cheap. For domestic marketers, this quarter’s experiment is one thing: carve 5% of the search-ad budget, put it into conversational, and compare conversion rates on the same basis.

transformers swallowed GGUF: the bottleneck moves to RAM

6. Transformers now runs llama.cpp quants

Hugging Face announced that GGUF now loads directly via from_pretrained. Speed is matched by reusing llama.cpp’s ggml kernels, and the post includes a table showing Qwen3.5 4B shrinking from BF16 8.42GB to 2.74GB at Q4_K_M. Top priorities: Apple Silicon and the Qwen3.5 architecture. You can even serve an OpenAI-compatible server with transformers serve.

From a hardware standpoint there is exactly one conclusion. The moment performance that llama.cpp monopolized merges with the standard Python ecosystem, local inference moves from lab experiment to mandatory developer equipment. The first bottleneck isn’t the GPU but unified memory capacity. Cases of running 27B-class models on 32GB MacBooks are already being cited. This lines up exactly with the view that DRAM spot prices are a leading indicator of the AI cycle. Even the 4090 I praised earlier ultimately comes back around to the next generation’s memory capacity problem.

So the recommendation is simple. In your next build, add a tier of DDR instead of a tier of VGA. The trend isn’t models getting smaller — it’s loading more onto the same memory.