Five independent axes
Knowledge and reasoning (can the model recall facts correctly and reason logically?), coding (does it solve programming tasks reliably?), agentic ability (can it carry out multi-step tasks autonomously using tools?), human preference (do people prefer its answers in a head-to-head comparison?), and cost efficiency (what does an answer cost relative to its quality?) – these are five largely independent axes. A model can be far ahead on one and clearly behind on another.
Why an announcement almost always shows just one axis
Making a new model the market leader on all five axes at once is rare. So an announcement almost always picks the axis where the new model performs best – that's not deception in a legal sense, but it's an incomplete statement when the other four axes go unmentioned.
How to take a claim apart
For any headline of the form "Model A beats Model B", the same follow-up question is worth asking: on exactly which of the five axes? And what do the other four show? A more nuanced picture often emerges – for example, a model that leads on coding but clearly trails a cheaper competitor on cost efficiency.
Why this matters to you as a decision-maker
Which axis matters to you depends on your own use case – a customer-service chatbot needs different strengths than a coding assistant. A vendor announcement that names only one axis almost never answers the question that actually matters: how does the model perform on the axis relevant to your specific use case?