Skip to content
Beyond Prompt AI Studio
Comparing AI models properly

Knowledge and reasoning benchmarks: MMLU, GPQA, and the saturation trap

"How AI performance is actually measured" introduced objectively-correct benchmarks as the first of three measurement paradigms. Their best-known examples – MMLU and GPQA – show a recurring pattern: a benchmark appears, becomes the standard, and then loses its power to differentiate within a few years.

Assessment as of: July 2026 – research moves fast, this assessment may change.

Tap a milestone to learn more.

Four examples – worth remembering

How MMLU and GPQA work

MMLU (Massive Multitask Language Understanding) consists of multiple-choice questions across 57 subjects, from law to medicine. GPQA (Graduate-Level Google-Proof Q&A) goes a step further: its biology, chemistry, and physics questions are written to be so demanding that even PhD holders outside the specific field only answer about 34% correctly – a deliberately high human baseline.

The saturation trap

MMLU now saturates above 88% for leading models – the benchmark barely differentiates top-tier models from each other anymore. GPQA Diamond, its hardest subset, long stayed well below that, but leading models now reach high scores there too and it's approaching saturation. A saturated benchmark barely answers "is Model A better than Model B?" anymore, because both sit close to the ceiling.

A pattern that keeps repeating

This has already happened once before: the language-understanding benchmark GLUE appeared in 2018 – within about a year, models had already surpassed the human baseline. In response, the deliberately harder successor SuperGLUE appeared in 2019. There too, the same thing quickly showed up: on most of its subtasks, models were already at or above human crowdworker level. The field then moved on to MMLU and later GPQA – following the same pattern now repeating with MMLU and GPQA themselves.

Why this matters to you as a decision-maker

A single MMLU or GPQA score comparing two top-tier models today says considerably less than it did just a few years ago – both often sit close together near the ceiling. Anyone judging a vendor claim based on a well-known, widely used benchmark should check whether that benchmark still differentiates the compared models at all, or whether it's already saturated.

Key takeaways

  • MMLU and GPQA are objectively-correct knowledge/reasoning benchmarks with multiple-choice questions and a known correct answer.
  • MMLU saturates above 88% for leading models; GPQA Diamond also reaches high scores and is approaching saturation.
  • A saturated benchmark barely differentiates top-tier models from each other anymore.
  • The same pattern already happened once: GLUE (2018) was surpassed within about a year, and its successor SuperGLUE (2019) saturated quickly too.
  • A vendor comparison based on a widely used benchmark should check whether it still differentiates the compared models, or is already saturated.

Quick check: did it land?

1 / 3

According to this module, why are GPQA questions deliberately written to be so hard?

Curious how meaningful common benchmarks still are?