Today’s model news is about how you measure the numbers: a new compression method that halves a 70B, an open dataset exposing the protocol behind the scores, and the changes that will actually reach my 4090 — I read the primary sources and did the math.
🌐 · English · 中文 · Español · 한국어 · · English Hub
How to cut 70B by 50% and keep MMLU 23 points — turning block removal into a physics problem
Post by Multiverse Computing (September 21) tackles the problem of erasing entire transformer blocks with a binary optimization model called an Ising glass. Unlike prior methods that slice out contiguous chunks, it runs only forward and backward passes on calibration data to rank combinations of ‘which blocks to keep.’ Key numbers: compressing Llama-3.3-70B by 50% beats prior SOTA block-removal methods by 23 points on MMLU, and even on models with heterogeneous block structure like Nemotron-3-Nano-30B-A3B it surpasses SOTA on AIME25 and GPQA.
My take: the per-4090 math is fun. Putting 35B — half of 70B — at 4-bit needs about 17.5GB in weights alone; subtract KV-cache headroom and it barely fits in 24GB — the first genuinely ‘true 70B-class’ path. But 50% compression isn’t a number you get without retraining, and benchmark scores don’t guarantee perceived quality. Still, it’s enough of a signal that ‘contiguous chunk cutting’ has effectively expired.
Benchmark scores are curves, not points — why UK AISI published the protocol
Released by UK AISI and the EvalEval Coalition (September 22)has released Evaluation Cards with verified results and configuration details for five benchmarks — HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0. Models covered include Claude Opus 4·4.5·4.6, GPT-5·5.2·5.4 and six in total. The highlight is the companion paper — Humanity’s Last Exam scores moved along curves with inference token budget and protocol, and with correct-answer feedback the models solved more problems the more tokens they spent.
My take: a leaderboard line is marketing the moment the condition ‘n tokens allowed’ is omitted. If the same model on the same benchmark scores differently just because the settings changed, the comparison itself needs to be redesigned. In my own local runs, the same model at temperature 0.7 versus 0.2 already feels different — the industry is only now starting to admit that in writing. When citing a benchmark, if there’s no settings link, it’s safer not to use that number.
7 hours saved per week and 1 billion people — what AI usage data tells us
OpenAIMentalHealthBench (September 23) was released. More than 80 licensed professionals across 22 countries took part; synthetic conversations are generated from real usage patterns at the scale of 1 billion weekly users and scored on 10 dimensions including safety, context awareness, and user autonomy. That same week GoogleNew ATLAS data reports that nearly half of scientists use AI daily and save about 7 hours a week, and that in the US computer and math occupations account for 30% of AI use — twice the global average.
My take: 7 hours is 17.5% of a 40-hour week — far more honest than ‘productivity revolution’ copy, and therefore a bigger number. Both announcements share one thing: they designed their own measurements instead of outsourcing them. Now that the scorecard is public, future model releases will have to be verified broken down along these same ten dimensions.
The bottleneck in browser inference was kernels, not the model — 207 WebGPU kernels released
Hugging Face@huggingface/kernels (September 1) released 207 WebGPU kernels. Each kernel is a standalone repo with a manifest, conformance tests, and benchmark cases, and a Fleet that grades GPUs directly in the browser collects performance data on real hardware the lab doesn’t cover. On the sameoMLX maintainer Jun Kim joinswas also announced — the first case of dedicated MLX-ecosystem staffing.
My take: a runtime can never beat its kernels. Splitting kernels into version-controlled Lego bricks is the engineering inflection point where browser inference moves from ‘demo’ to ‘infrastructure’ — and with a dedicated Apple Silicon team now, my M5 Mac will soon have a job as a spare inference node. If local setup is unfamiliar, myLocal installation guideis the starting point.
Frequently asked questions
Can I download the block-removed compressed models right now?
The post is methodology- and benchmark-focused, so a released finished checkpoint isn’t confirmed. But since it works with a calibration pass alone, this is the kind of research the community will reproduce quickly.
Will a 70B compressed 50% run on an RTX 4090?
35B at 4-bit is about 17.5GB, so it theoretically fits in 24GB. But to leave room for KV cache and context length, you need measurements at 32K or below, and minimum VRAM per use case isMy reference table— check it via
If a model scores low on MentalHealthBench, should I avoid it for counseling?
OpenAI itself draws the line: ChatGPT is not a substitute for therapy. This benchmark measures the overall quality of everyday conversation, not crisis-response avoidance, so scores are a reference metric, not a safety guarantee.
Sources: Hugging Face official blog, OpenAI official announcement, Google AI & Economy ATLAS. Figures are as announced and may change with subsequent replication results.