Comparing AI models properly
Learning goal
How to seriously measure model performance, why benchmark numbers often say less than they claim, and how to make sense of vendor comparisons yourself.
01
How AI performance is actually measured: the building blocks you need
Start module →
02
Why "Model A beats Model B" is almost never the whole story
Start module →
03
Knowledge and reasoning benchmarks: MMLU, GPQA, and the saturation trap
Start module →
04
When people judge: Chatbot Arena, Elo ratings, and their blind spots
Start module →
05
The judge is just a model too: when AI scores AI
Start module →
06
Why good test scores can deceive: contamination and the training-data curse
Start module →
07
The practical test: agentic benchmarks, time horizons, and how to compare models yourself
Start module →