GPU 시세 · VRAM · 온디바이스 AI

MLPerf Inference v6.1 results and hands-on reviews of Apple’s new hardware landed in the same week. What the two share isn’t a core-count race but ‘the economics of the whole box.’ I read the originals in full and pulled out only the numbers that can be computed.

🌐 · English · 中文 · Español · 한국어 · · English Hub

Vera Rubin NVL72’s 3.7x isn’t a chip score — where the number actually comes from

NVIDIA official announcement (September 16)According to the report, in the MLPerf Inference v6.1 preview the Vera Rubin NVL72 posted up to 3.7× the throughput of the GB300 NVL72 on the Qwen3-VL benchmark (vLLM + Dynamo config) and 2.5× on DeepSeek-R1 (TensorRT-LLM). But the numbers I paid more attention to are elsewhere: 99% linear scaling efficiency on a 4-rack GB300, 288-GPU configuration, and the fact that 1.6× over v6.0 came purely from software optimization within six months.

My take: reading 3.7x as ‘chip performance 3.7x’ is wrong. The 72-GPU configuration stays the same; the difference comes from the NVLink domain and scheduling. 99% scaling means adding racks increases throughput almost proportionally — quantitative proof that the interconnect, not chip clocks, is the bottleneck. It also catches my eye that the benchmark models are DeepSeek-R1 and Qwen3-VL — the same family I run on a single 4090, so the vLLM optimizations polished above will eventually reach the config file on my desk.

The M5 Ultra Mac Studio 256GB is $48 per GB — an eighth of an H100’s, with conditions

IT World Korea reviewThe top configuration tested runs 36-core CPU, 80-core GPU, 256GB memory, 4TB storage for 12,299 dollars. Geekbench 7 multicore 52,350, Metal GPU 360,019, storage write 13,941MB/s — 2x the previous generation — and AI performance up to 4.3x the M3 Ultra (Apple’s claimed figures). The reviewer ran it headless, loaded models with LM Studio, and ran it as a render/inference server at roughly 10 cents an hour based on ~480W peak.

My take: do the math yourself — $12,299 divided by 256GB is about $48 per GB. That’s 1/8 of the $375 per GB you get pricing an H100 80GB at $30,000, and cheaper than the $67 per GB of a 4090 24GB (~$1,600). But memory bandwidth is a server blowout, so token generation speed suffers. Hence the split conclusion — for an ‘engine room’ role quietly keeping a heavily quantized model always on standby, it’s the best cost-performance available; not for real-time serving. For minimum VRAM per use case, myVRAM reference tablefor the contrast.

Why iOS 27 hyped the new models but got faster on old iPhones — second-by-second measurements

iPhone 11 Pro Max measuredCamera launch went from 0.98s to 0.65s, post-capture thumbnail check from 3.61s to 2.44s, and Maps launch from 3.01s to 1.36s — 2.2x faster. Thermals reportedly improved too. But Siri AI and Apple Intelligence are exclusive to iPhone 15 Pro and later.

My take: the paradox where AI features land on new models while the felt speedup lands on old ones. It’s the result of animation-pipeline and background-management tuning — and in AI marketing’s shadow, old devices effectively gain 1~2 more years of life. The used market feels it too: iPhone 12/13 listings that run iOS 27 smoothly will hold perceived value longer and their price curve may flatten, and unlike a training-GPU budget, a phone upgrade is a budget you can postpone.

The trust anchor for authentic photos has moved from metadata to the camera sensor

Apple reference-image whitepaperTo sum up. The iPhone 18 Pro signs the pixels with a cryptographic signature via the Secure Enclave and Private Cloud Compute the instant you press the shutter, sealing immutable metadata and its own timestamp. Apple’s stated reason for calling the industry-standard C2PA weak: signatures are attached after capture and link the photographer’s identity to the image — for photographers in conflict zones, that structure paints a target on them. But: main camera only, no retroactive application, China unsupported at early launch, view-only in Europe.

My take: the essence is the trust anchor shifting from ‘after-the-fact proof’ to ‘signature at the moment of capture.’ At the same time, the problem remains that there is only one anchor — Apple. In a market where a large share of news photos are shot on iPhones, the right to judge what’s ‘real’ now belongs to one company, and authenticity verification itself is starting to fracture along geopolitics, as with the China/Europe split. How fast the C2PA camp catches up with sensor-level signing is the first thing to watch.

Frequently asked questions

What does Vera Rubin’s 3.7x have to do with local users?

Not directly comparable, but the benchmark models (DeepSeek-R1, Qwen3-VL) and the inference framework (vLLM) are the same family as local setups. Batch and cache optimization techniques validated at rack scale transfer straight down to a desk-side configuration.

Can I run 70B-class models on an M5 Ultra 256GB?

A 70B-class 4-bit quantization weighs about 40GB, so several can stay loaded at once with room to spare. The tradeoff: bandwidth limits mean generation speed can trail desktop GPUs, so it’s ideal for low-frequency, always-on standby duty.

Do reference images work on older iPhones too?

No. It’s exclusive to the iPhone 18 Pro’s main camera, with no backporting, and initially it’s unsupported in China and view-only in Europe.

Source: NVIDIA official blog MLPerf v6.1 results (September 16), IT World Korea measured review, Apple reference image whitepaper. Prices and figures are as announced and may change.