AI · session log · 02:14 BST

The half-life of a benchmark.

MMLU was a real measurement in 2023 and a polite formality by 2025. HumanEval, GSM8K, MATH all went the same way. A benchmark is alive only as long as the gap between models still maps to something a customer can feel; once the labs train against it, the rate card on the wall becomes decoration.

studio

mirpur · road 7

console

SSL 4000-class

session

demo grading

scale

in-house, 2003

meter

VU peak +3

status

still posted

scroll · faders down
Sajjad bhai keeps the demos in a tin biscuit box on the second shelf above the patchbay, and the box has been there since the studio opened on Road 7 in Mirpur in 1998. Tonight the box has eleven cassettes, each with a name someone wrote in biro across the J card by whichever lyricist dropped it off, and the session clock above the talkback panel reads 02:14. The way he grades them hasn't changed since 2003, when he and his older brother sat at this same console one Friday evening with a bottle of Chivas they couldn't quite afford, and worked out their own scale of ten points across six axes. Hook, mix, vocal, lyrical, arrangement, and a sixth column the brother stubbornly called "Dhaka factor" and never quite defined. For about a decade the scale was uncannily good at predicting which demos would chart on Radio Foorti and which would die on the cutting room floor.
Tonight the scale isn't predicting anything. The two demos he's rated highest, an 8.6 and an 8.4, are slow ghazal pop of the kind he grew up grading near the top of the list. Neither will survive the first week of any TikTok the artists' cousins record next month. The two he's rated lowest, a 4.2 and a 3.9, both glitchy things where the bass sits up front and he can't quite hear properly through the NS-10s, are exactly the shape of song that will have ten million streams by August. The scale itself hasn't gone wrong in any technical sense. The audience around it has simply moved, and around 2017, on his own private accounting, the rate card on the wall started measuring something other than what it was built to measure. It took him another four years of grading sheets in the same column to notice.
This is also, to a degree he finds slightly embarrassing, the story of every AI benchmark you've heard of.
solo
ch · 14
monitor · isolated track
A benchmark stays alive only as long as the gap between two models on it still maps to something the user actually feels in the chair. The moment that gap becomes a polite formality, the whole apparatus turns into the rate card by Sajjad bhai's door.
MMLU is the cleanest example, and the one I keep coming back to because we leaned on it heavily inside an internal agentic evaluation rig for most of 2024. Hendrycks and the paper's other authors published MMLU in 2020 as a set of 57 tasks spanning US history to college medicine, and for two years it was the most honest single number you could quote about a frontier model. By the back end of 2023 the gap between leading models had narrowed to roughly two percentage points, almost all of it inside the confidence interval of the benchmark itself, and by the middle of 2024 every credible model was sitting north of 86%. Stanford's CRFM team had been running HELM the whole time as a kind of standing committee on this exact problem. You can read their quarterly retros as a slow obituary for a generation of leaderboards built around one number. Saturation at that point on the curve does not mean the models have stopped improving. It means the meter has stopped registering whatever improvement is still happening underneath.
The same shape played out faster on the coding benchmarks. HumanEval, the original Python set of 164 problems OpenAI shipped with Codex, was a real instrument in 2021, an interesting one in 2022, and decoration by the time GPT-4 was passing it cold in early 2023. GSM8K, a maths set aimed at grade school students, lasted slightly longer because the chain-of-thought prompting era stretched its useful life. By the middle of 2024 every serious lab was sitting in the high nineties, and the labels on the leaderboard had started reading like a Bangladeshi A-level result sheet, where the only meaningful axis is who finished in the absolute top tier and the rank between them is mostly noise. MATH and ARC went the same way over 2024 to 2025. Big-Bench Hard, which was supposed to last longer because it was deliberately designed to be hard, made it about eighteen months before the same flatlining set in.
The pattern is consistent enough that you can almost set your watch by it. A benchmark gets published. The labs train against it, sometimes deliberately and sometimes via the slower osmosis of the test set turning up somewhere inside data scraped from the whole internet. The scores climb steeply for a year or two, then asymptote near the ceiling, and then everyone is north of 90% and the only honest thing to do is admit the meter has stopped. Rich Sutton's Bitter Lesson sits in the background of all of this. His point, that general methods riding on more compute eventually crush approaches tuned by hand, is also the reason every benchmark with a fixed surface area has the half-life it does. The compute keeps coming, the surface area doesn't grow, and the curve flattens.
Spotting one going dead before everyone else does is not, in my experience, a numerical exercise. It is a matter of paying attention to two things the leaderboard itself will not tell you. The first is whether the benchmark still produces ranks that surprise the practitioners. If you ask three serious engineers which model is best at the thing the benchmark claims to measure, and all three guess the same model and are right, the benchmark has become a restatement of what the field already knows. The second is whether the gap on the metric still corresponds to a gap on the customer-facing task. Ethan Mollick has been writing about this for two years in his Substack, and his most useful framing, paraphrased loosely, is that the only honest test of an eval is whether the model that wins it also wins the blind A/B with the user.
What replaces it is, awkwardly, a moving target. Hamel Husain's argument that the only honest unit of progress in an LLM application is your own internal eval, kept against your own product, is the version of this I've come to agree with. The labs are the ones who get to build the model. What's left for the rest of us in the application layer is to build the meter that grades it on the work the customer actually does, with our retrieval pipeline and our prompts and our particular edge cases. It will go stale eventually too, on the same curve, and you'll catch it going stale because the engineers in the next standup will start agreeing about which version was better before the metric does. That agreement, the moment three people in a room can predict the score, is the half-life expiring in real time.
Sajjad bhai told me last winter, between two takes on a fusion track he was producing for a singer who had come down from Sylhet for the week, that he was leaving the rate card up on the wall anyway. The card is from 1998, and the numbers on it stopped meaning anything around 2017, but the wall looks wrong without it, and he wants every younger engineer who comes through the studio in the next decade to understand that whatever scale they are using on a Monday morning is going to look exactly that quaint by the time somebody else is sitting at the console. Some of the meters in front of you are still measuring what they claim to measure, and paying attention to which ones have quietly turned into wall decoration is a job that never really finishes, the same way that tin biscuit box of cassettes has stayed on the second shelf above the patchbay since 1998 because nobody has found a reason to move it either.