GPU 시세 · VRAM · 온디바이스 AI

From cases where Codex drove revenue to ChatGPT ads expanding into Asia — today’s six picks from a developer’s perspective. No bare link-drops: I read the originals and added each story’s key numbers and my take.

🌐 · English · 中文 · Español · 한국어 · · English Hub

1. Codex built a sales rep’s demo — 40–60 engineer-hours freed up

Proaction boosts sales 60% and saves 75+ hours with Codex

This is the story of Proaction, a vehicle fleet management software startup. Their sales COO isn’t a developer — she throws call recordings, emails, and spreadsheets of prospective customers into Codex andInteractive demo screen with that company’s vehicle datathemselves in 30–45 minutes. 4–6 a month. If an engineer built them it’d take 10 hours each, so 40–60 hours a month are saved, and the rate of deals moving from first contact to actual solution-development stage rose 50–60%.

My take: what’s being priced isn’t model performance butThe bottleneck movesis the point. Demo production moved from the engineering queue to a salesperson’s on-the-spot task — if demos were your bottleneck, this workflow breakdown is worth porting as-is, more than any spec table.

2. Harvey raises the context ceiling for legal drafts

Harvey turns legal context into stronger drafts with GPT-6 Astra

Legal AI Harvey showed a flow that pulls the full matter context — case law, counsel documents, filings — into drafting, and pins per-lawyer preferences (numbered lists, source priority, color by priority) in a memory panel. Document formatting and context incorporation are reportedly markedly better than previous models.

My take: it’s no accident that law took the front row — large input document volume and strict output formats meanContext length and structuring equal qualitybecause it’s a domain that becomes exactly that. If you’re running RAG locallyVRAM reference tableand compute your cache budget first.

3. invideo — 3× color-grading success rate, 50 custom effects a day

How invideo improves color grading 3x with GPT-6 Astra

Report from agentic video editor invideo: separation tasks — change the background, keep subject tones — demanded frame tracking and used to fail often; the success rate is now roughly 3x. They built about 50 custom effects per day from descriptions and references, and complex editsFewer reasoning steps and fewer output tokensis what I plan for.

My take: the substance of agent cost is the number of reasoning steps. Fewer output tokens on the same task means latency and the bill both shrink — if you’re picking an agent, look at whether it publishes step counts rather than benchmark scores.

4. ChatGPT ads, 7 more markets in Southeast Asia and Taiwan — 1 billion dollars annualized within 200 days of launch

ChatGPT Ads expands to Southeast Asia and Taiwan

Rollout started in Indonesia, Malaysia, the Philippines, Singapore, Thailand, Vietnam, and Taiwan — over 60 countries cumulative. One number: since the ads launched1 billion dollars annualized revenue in under 200 daysachievement speed. It was reconfirmed that ads appear only on the free and low-tier plans, premium subscriptions stay ad-free, and conversation content is never disclosed to advertisers. Korea was already part of the previous rollout wave.

My take: the value of an ad inventory comes from purchase-consultation intent — the same reason search ads are expensive. For Korean marketers, the format has already landed in Korea, so it’s time to check what consultation phrases competitors are locking down.

5. Airbnb expands GPT-6 Astra across engineering — doc work from 20 rounds to 3~4

Airbnb widens access to GPT-6 Astra and OpenAI frontier models

Airbnb, which has used the GPT-5.6 family for coding agents, expanded Astra access across its entire product and engineering organization via the API and Amazon Bedrock paths. Interesting data point: for a single user doing strategy documents rather than coding, output reachedFrom 20+ rounds on other models to 3–4 roundshas been shortened.

My take: at scale, the winner is decided by rounds. It’s not one person saving an hour but a queue problem for the whole organization — which is why budgets now pass even in teams that used to be uncooperative.

6. Transformers now runs llama.cpp quantizations as-is — GGUF via from_pretrained

Transformers now runs llama.cpp quants

Hugging Face Transformers now loads GGUF directly. Reusing the ggml kernels that power Ollama and LM Studio through the kernels library, it gets close to llama.cpp performance; the initial optimization targets areApple Silicon and the Qwen3.5 architectureis the answer. Example table: Qwen3.5-4B from BF16 8.42GB to Q4_K_M 2.74GB — we recommend Q4_K_M as the starting point.

My take: GGUF has operated as the de facto standard for local execution, yet it was cut off from the Transformers ecosystem (datasets, runners, evaluation). Once that wall falls, the quantized models waiting in folders come into the pipeline. For Mac users, this is a story starting today.

Frequently asked questions

What should you prepare first for demo automation?

The Codex environment itself is the precondition, so setting up call-recording, email, and CRM integrations matters before the model itself. If you want a local alternative, automate context collection and leave generation toLocal coder guide— go with this configuration.

GGUF support — worth switching to instead of Ollama?

For generation alone, Ollama is still the easy path. When you need to plug GGUF into fine-tuning, evaluation, and dataset pipelines, Transformers just opened that road.

Are ChatGPT ads already in Korea too?

Yes — after Australia, New Zealand, Japan, South Korea, and India, 7 Southeast Asian markets were added. Only free/low-tier plan accounts are eligible for exposure.

Source: manufacturer official announcements and published figures. Prices and specs can change — verify before deciding.