How MMLU and GPQA work
MMLU (Massive Multitask Language Understanding) consists of multiple-choice questions across 57 subjects, from law to medicine. GPQA (Graduate-Level Google-Proof Q&A) goes a step further: its biology, chemistry, and physics questions are written to be so demanding that even PhD holders outside the specific field only answer about 34% correctly – a deliberately high human baseline.
The saturation trap
MMLU now saturates above 88% for leading models – the benchmark barely differentiates top-tier models from each other anymore. GPQA Diamond, its hardest subset, long stayed well below that, but leading models now reach high scores there too and it's approaching saturation. A saturated benchmark barely answers "is Model A better than Model B?" anymore, because both sit close to the ceiling.
A pattern that keeps repeating
This has already happened once before: the language-understanding benchmark GLUE appeared in 2018 – within about a year, models had already surpassed the human baseline. In response, the deliberately harder successor SuperGLUE appeared in 2019. There too, the same thing quickly showed up: on most of its subtasks, models were already at or above human crowdworker level. The field then moved on to MMLU and later GPQA – following the same pattern now repeating with MMLU and GPQA themselves.
Why this matters to you as a decision-maker
A single MMLU or GPQA score comparing two top-tier models today says considerably less than it did just a few years ago – both often sit close together near the ceiling. Anyone judging a vendor claim based on a well-known, widely used benchmark should check whether that benchmark still differentiates the compared models at all, or whether it's already saturated.