Skip to content
Beyond Prompt AI Studio

Comparisons

Closed-source LLMs compared

Six proprietary LLM APIs, scored uniformly on the same seven criteria - with a source link and date for every entry, so you can verify the numbers yourself. This category is aimed at developers and technical decision-makers building an application on an API - not end users of a finished chat app (see our "AI Chat Assistants" category for that). Switch between the individual and business perspective at the top of the table. At the bottom you'll find scenario recommendations instead of a single "winner": which API fits depends on your specific situation.

A comparison of closed-source LLM APIs weighs offerings like GPT, Claude, Gemini, Grok, Cohere and Kimi K3 from a developer's perspective - from price per token to rate limits to enterprise SLAs - for anyone building their own application on an API.

How we score

At a glance

Value for moneyPrivacy &hostingPerformance &context windowRate limits &availabilityEcosystem &toolingMultimodalityEnterprisereadiness &compliance

Calculated automatically from the scores below - the provider highlighted in green has the highest average score across all criteria, not an editorial pick. Toggle providers in the legend on or off.

View for
Scores stay the same - only strengths, weaknesses, and the verdict adapt to the perspective.

Last data review: 08/17/2026

Click a row for strengths, weaknesses, and our take.

Which tool fits you?

The highest raw-performance benchmark (coding/analysis)

GPT (API)

The highest benchmark score among all six APIs compared (96.2% SWE-bench) at an unchanged flagship price.

Building agents and tools with MCP

Claude (API)

The origin and deepest integration of the MCP standard, plus one of the highest benchmark scores in this comparison.

Multimodal applications from a single provider

Gemini (API)

The broadest multimodality (image, video, audio, four image-generation tiers) and the best certification standing.

Maximum cost efficiency outside the EU

Grok (API)

Still the cheapest flagship pricing tier among the APIs - for EU customers, currently only the predecessor Grok 4.3 is usable.

A cheap, high-performance API for non-critical data

Kimi K3

The third-highest benchmark score in this comparison at a low cost - but without certifications or EU hosting, so suitable only for non-critical data.

Key takeaways

  • GPT (API) scores 96.2% on SWE-bench with GPT-5.6 Sol at an unchanged flagship price of $5/$30 per million tokens - though the top of the leaderboard now belongs to Claude Opus 5 at 97.0%.
  • Claude (API) fields the strongest model in this comparison with Opus 5 (97.0%) - and at half the price of its own Fable 5 ($5/$25 vs. $10/$50). Anthropic recommends Opus 5 as the default starting point; Fable 5 remains the special case for maximum capability, but without a zero-data-retention option.
  • Kimi K3 (Moonshot AI) is new in this comparison: the third-highest benchmark score (93.4%) at a fraction of the cost - but with no certifications, EU hosting, or ZDR option at all, and as of today still without open weights (promised by Moonshot for 2026-07-27).
  • Grok (API) made the biggest benchmark jump of any vendor with 4.5, but per xAI itself it isn't yet available in the EU - EU customers should stick with the predecessor Grok 4.3 for now.
  • Cohere surprisingly pivoted to open weights with Command A+ (Apache 2.0, free) - but raw performance remains the weakest in this comparison.
  • Gemini (API) still offers the broadest multimodality and the best certification coverage among the six APIs.

Self-host, use an API, or buy a ready-made solution? Build vs. buy vs. API

Find the right approach for your own data: RAG-or-Fine-Tuning Finder

Frequently asked questions

Which LLM API is best for coding applications?

GPT (API) now scores the highest benchmark among the six APIs compared here with GPT-5.6 Sol (96.2% SWE-bench) - just ahead of Claude's new Fable 5 (95.0%) and Kimi K3 (93.4%). Claude, though, remains the origin of the MCP (Model Context Protocol) standard for tool integrations and still has the deepest ecosystem there.

Which LLM API is the cheapest?

Grok (API) still has the cheapest flagship pricing per million tokens among the APIs with a full enterprise offering - though for EU customers, only the predecessor Grok 4.3 is currently available. Kimi K3 undercuts Grok on cache-hit input pricing, but comes in above it on output pricing.

Is there an LLM API with on-premise deployment for regulated industries?

Cohere is still the only option among the six APIs compared with true private/on-premise deployment (Model Vault, BYOC) - since May 2026 also with open weights (Command A+, Apache 2.0) for anyone who wants to self-host.

Is Kimi K3 an open or a closed model?

As of 2026-07-21, Kimi K3 is a pure API model with no publicly available weights - Moonshot AI has announced an open-weight release for 2026-07-27, but hasn't delivered it yet. That's why Kimi K3 currently sits in this category rather than our open-source LLMs category; if the announcement lands, we'll reassess where it belongs.

Why is Cohere listed here even though Command A+ has open weights?

Cohere is evaluated in this category primarily for its API/enterprise business model (Model Vault, managed Rerank/Embed services, a sales process for enterprise terms), not for the weights themselves. Unlike the models in our open-source category, Cohere's managed platform is the focus here - Command A+'s open weights are an additional self-hosting option, not a business-model pivot.

What's the difference between this category and "AI Chat Assistants"?

This category evaluates APIs from a developer's perspective (price per token, context window, SDKs, rate limits) for building custom applications. The "AI Chat Assistants" category instead compares the finished consumer chat apps (ChatGPT, Claude.ai, the Gemini app, etc.) for end users.