Leaderboard · 2026

AI Model Leaderboard — 2026

Every model you can call through AIgateway, ranked by the benchmarks engineers actually care about: SWE-bench for agentic coding, MMLU + GPQA for reasoning, HumanEval for code generation — plus input and output price per 1M tokens and context window size. Numbers are sourced from each provider's reported eval results; blank cells mean the lab has not published that benchmark. Sort any column. The same data is available as JSON at /api/leaderboard.

1031 models across 82 providers. Every row is a click away from a full model page with pricing, quickstart code, and a live playground.

133 of 1031 models · JSON
ModelProviderSWE-benchMMLUGPQAHumanEvalInput $/MOutput $/MContext
Grok 4.6
xai/grok-4.6
xAI$1.00$1.00500K
Claude Opus 5
anthropic/claude-opus-5
Anthropic$5.00$25.001M
Qwen3.8 Max
alibaba/qwen3.8-max
Alibaba$2.00$6.001M
DeepSeek V4 Pro
deepseek/deepseek-v4-pro
Deepseek$2.40$4.80131K
Kimi K3
moonshot/kimi-k3
Moonshot$3.00$15.001.0M
GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI$5.00$30.001.1M
Claude Fable 5
anthropic/claude-fable-5
Anthropic$10.00$50.001M
Gemini 3.6 Flash
google/gemini-3.6-flash
Google$1.50$7.501.0M
GLM-5.2
zai-org/glm-5.2
Zai-org$1.40$4.40262K
Gemini 3.1 Pro
google/gemini-3.1-pro
Google89.6$2.00$12.001M
MiniMax M3
minimax/m3
MiniMax$0.30$1.201M
Kimi-K2.6
moonshot/kimi-k2.6
Moonshot68.292.7$0.95$4.00262K
O3
openai/o3
OpenAI$2.00$8.00200K
Gemini 2.5 Pro
google/gemini-2.5-pro
Google$1.25$10.001M
GPT-5.5 Pro
openai/gpt-5.5-pro
OpenAI$30.00$180.001M
GLM-5.1
zai-org/glm-5.1
Z.ai$1.40$4.40200K
Gemini 3.5 Flash-Lite
google/gemini-3.5-flash-lite
Google$0.30$2.501.0M
Gemini 2.5 Flash
google/gemini-2.5-flash
Google$0.30$2.501M
Gpt-Oss-120b
openai/gpt-oss-120b
OpenAI$0.35$0.75128K
Sonar Reasoning Pro
perplexity/sonar-reasoning-pro
Perplexity$2.00$8.00127K
Nemotron-3-120b-A12b
nvidia/nemotron-3-120b-a12b
Nvidia$0.50$1.50256K
GPT-5.4 Mini
openai/gpt-5.4-mini
OpenAI84.588.6$0.75$4.50128K
Sonar Deep Research
perplexity/sonar-deep-research
Perplexity$2.00$8.00127K
Gemma-2b-IT-Lora
google/gemma-2b-it-lora
Google$0.030$0.0608K
Llama-3.2-3b-Instruct
meta/llama-3.2-3b-instruct
Meta$0.051$0.3480K
Mistral-7b-Instruct-V0.2-Lora
mistral/mistral-7b-instruct-v0.2-lora
Mistral$0.050$0.1015K
Kimi-K2.7-Code
moonshot/kimi-k2.7-code
Moonshot$0.95$4.00262K
Llama-3.1-8b-Instruct-Fp8
meta/llama-3.1-8b-instruct-fp8
Meta$0.15$0.2932K
Llama-3.2-1b-Instruct
meta/llama-3.2-1b-instruct
Meta$0.027$0.2060K
Glm-4.7-Flash
zai-org/glm-4.7-flash
Zai-org$0.060$0.40131K
Llama-2-7b-Chat-HF-Lora
meta-llama/llama-2-7b-chat-hf-lora
Meta-llama$0.040$0.0808K
Llama-3.3-70b-Instruct-Fp8-Fast
meta/llama-3.3-70b-instruct-fp8-fast
Meta$0.29$2.2524K
Granite-4.0-H-Micro
ibm-granite/granite-4.0-h-micro
Ibm-granite$0.017$0.11131K
Qwen2.5-Coder-32b-Instruct
qwen/qwen2.5-coder-32b-instruct
Alibaba Qwen$0.66$1.0033K
Gemma-Sea-Lion-V4-27b-IT
aisingapore/gemma-sea-lion-v4-27b-it
AI Singapore$0.35$0.56128K
Qwen3-30b-A3b-Fp8
qwen/qwen3-30b-a3b-fp8
Alibaba Qwen$0.051$0.3433K
Gemma-7b-IT-Lora
google/gemma-7b-it-lora
Google$0.080$0.164K
Gemma-4-26b-A4b-IT
google/gemma-4-26b-a4b-it
Google$0.10$0.30256K
Mistral-Small-3.1-24b-Instruct
mistralai/mistral-small-3.1-24b-instruct
Mistral$0.35$0.55128K
Gpt-Oss-20b
openai/gpt-oss-20b
OpenAI$0.20$0.30128K
Llama-4-Scout-17b-16e-Instruct
meta/llama-4-scout-17b-16e-instruct
Meta$0.27$0.85131K
Grok 4 Fast
xai/grok-4-fast
xAI$0.50$2.00256K
Grok 4
xai/grok-4
xAI$5.00$15.00256K
Claude Haiku 4.5
anthropic/claude-haiku-4.5
Anthropic80.185.2$1.00$5.00200K
Claude Opus 4.5
anthropic/claude-opus-4.5
Anthropic$5.00$25.00200K
Claude Opus 4.6
anthropic/claude-opus-4.6
Anthropic$5.00$25.001M
Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic72.590.495.1$5.00$25.001M
Claude Opus 4.8
anthropic/claude-opus-4.8
Anthropic$5.00$25.001M
Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic$3.00$15.00200K
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic62.187.293.4$3.00$15.00200K
Claude Sonnet 5
anthropic/claude-sonnet-5
Anthropic$2.00$10.001M
GPT-4.1
openai/gpt-4.1
OpenAI$2.00$8.001.0M
GPT-4.1 Mini
openai/gpt-4.1-mini
OpenAI$0.40$1.601.0M
GPT-4.1 Nano
openai/gpt-4.1-nano
OpenAI$0.10$0.401M
GPT-4o
openai/gpt-4o
OpenAI$1.25$5.00128K
GPT-4o Mini
openai/gpt-4o-mini
OpenAI$0.075$0.30128K
GPT-5
openai/gpt-5
OpenAI$1.25$10.00128K
GPT-5 Mini
openai/gpt-5-mini
OpenAI$0.25$2.00128K
GPT-5 Nano
openai/gpt-5-nano
OpenAI$0.050$0.40128K
GPT-5.1
openai/gpt-5.1
OpenAI$1.25$10.00128K
GPT-5.4
openai/gpt-5.4
OpenAI91.894.0$2.50$15.001M
GPT-5.4 Nano
openai/gpt-5.4-nano
OpenAI$0.20$1.25128K
GPT-5.4 Pro
openai/gpt-5.4-pro
OpenAI$30.00$180.001M
GPT-5.5
openai/gpt-5.5
OpenAI$5.00$30.001M
GPT-5.6 Luna
openai/gpt-5.6-luna
OpenAI$0.20$1.201.1M
GPT-5.6 Terra
openai/gpt-5.6-terra
OpenAI$2.00$12.001.1M
Gemini 2.5 Flash Lite
google/gemini-2.5-flash-lite
Google$0.10$0.401M
Gemini 3 Flash
google/gemini-3-flash
Google82.3$0.50$3.001M
Gemini 3.1 Flash Lite
google/gemini-3.1-flash-lite
Google$0.25$1.501M
Gemini 3.5 Flash
google/gemini-3.5-flash
Google$1.50$9.001.0M
Grok 4.20 Multi-Agent
xai/grok-4.20-multi-agent-0309
xAI$2.00$6.002M
Grok 4.20 Non-Reasoning
xai/grok-4.20-0309-non-reasoning
xAI$2.00$6.002M
Grok 4.20 Reasoning
xai/grok-4.20-0309-reasoning
xAI$2.00$6.002M
Grok 4.3
xai/grok-4.3
xAI$1.25$2.501M
Grok 4.5
xai/grok-4.5
xAI$1.00$1.00500K
MiniMax M2.7
minimax/m2.7
MiniMax$0.30$1.20128K
Qwen 3 Max
alibaba/qwen3-max
Alibaba$1.20$6.00256K
Qwen 3.5 397B A17B
alibaba/qwen3.5-397b-a17b
Alibaba$0.60$3.60256K
Qwen3.7 Max
alibaba/qwen3.7-max
Alibaba$2.50$7.501M
Qwen3.7 Plus
alibaba/qwen3.7-plus
Alibaba$0.40$1.601M
O3-Mini
openai/o3-mini
OpenAI$1.10$4.40200K
O4-Mini
openai/o4-mini
OpenAI$1.10$4.40200K
Qwen Flash
alibaba/qwen-flash
Alibaba$0.050$0.401M
Qwen Flash Character
alibaba/qwen-flash-character
Alibaba$0.050$0.40
Qwen Max
alibaba/qwen-max
Alibaba$1.60$6.40
Qwen-Omni Turbo
alibaba/qwen-omni-turbo
Alibaba$0.070$0.27
Qwen Plus
alibaba/qwen-plus
Alibaba$0.40$4.001M
Qwen Plus Character
alibaba/qwen-plus-character
Alibaba$0.50$1.40
Qwen Plus Latest
alibaba/qwen-plus-latest
Alibaba$0.40$4.001M
Qwen Turbo
alibaba/qwen-turbo
Alibaba$0.050$0.50
Qwen-VL Max
alibaba/qwen-vl-max
Alibaba$0.80$3.20
Qwen-VL OCR
alibaba/qwen-vl-ocr
Alibaba$0.070$0.16
Qwen-VL Plus
alibaba/qwen-vl-plus
Alibaba$0.21$0.63
Qwen3 14B
alibaba/qwen3-14b
Alibaba$0.35$4.20
Qwen3 235B A22B
alibaba/qwen3-235b-a22b
Alibaba$0.70$8.40
Qwen3 235B A22B Instruct
alibaba/qwen3-235b-a22b-instruct-2507
Alibaba$0.23$0.92
Qwen3 235B A22B Thinking
alibaba/qwen3-235b-a22b-thinking-2507
Alibaba$0.23$2.30
Qwen3 30B A3B
alibaba/qwen3-30b-a3b
Alibaba$0.20$2.40
Qwen3 30B A3B Instruct
alibaba/qwen3-30b-a3b-instruct-2507
Alibaba$0.20$0.80
Qwen3 30B A3B Thinking
alibaba/qwen3-30b-a3b-thinking-2507
Alibaba$0.20$2.40
Qwen3 32B
alibaba/qwen3-32b
Alibaba$0.16$0.64
Qwen3 8B
alibaba/qwen3-8b
Alibaba$0.18$2.10
Qwen3-Coder 480B A35B Instruct
alibaba/qwen3-coder-480b-a35b-instruct
Alibaba$1.50$7.50200K
Qwen3-Coder Flash
alibaba/qwen3-coder-flash
Alibaba$0.30$1.501M
Qwen3-Coder Next
alibaba/qwen3-coder-next
Alibaba$0.30$1.50256K
Qwen3-Coder Plus
alibaba/qwen3-coder-plus
Alibaba$1.00$5.001M
Qwen 3 Max Preview
alibaba/qwen3-max-preview
Alibaba$1.20$6.00256K
Qwen3-Next 80B A3B Instruct
alibaba/qwen3-next-80b-a3b-instruct
Alibaba$0.15$1.20
Qwen3-Next 80B A3B Thinking
alibaba/qwen3-next-80b-a3b-thinking
Alibaba$0.15$1.20
Qwen3-Omni Flash
alibaba/qwen3-omni-flash
Alibaba$0.43$1.66
Qwen3-VL 235B A22B Instruct
alibaba/qwen3-vl-235b-a22b-instruct
Alibaba$0.40$1.60
Qwen3-VL 235B A22B Thinking
alibaba/qwen3-vl-235b-a22b-thinking
Alibaba$0.40$4.00
Qwen3-VL Flash
alibaba/qwen3-vl-flash
Alibaba$0.050$0.40256K
Qwen3-VL Plus
alibaba/qwen3-vl-plus
Alibaba$0.20$1.60256K
Qwen3.5 122B A10B
alibaba/qwen3.5-122b-a10b
Alibaba$0.40$3.20256K
Qwen3.5 27B
alibaba/qwen3.5-27b
Alibaba$0.30$2.40256K
Qwen3.5 35B A3B
alibaba/qwen3.5-35b-a3b
Alibaba$0.25$2.00256K
Qwen3.5 Flash
alibaba/qwen3.5-flash
Alibaba$0.10$0.401M
Qwen3.5-Omni Flash
alibaba/qwen3.5-omni-flash
Alibaba$0.40$2.20
Qwen3.5-Omni Plus
alibaba/qwen3.5-omni-plus
Alibaba$1.40$8.30
Qwen3.5 Plus
alibaba/qwen3.5-plus
Alibaba$0.40$2.401M
Qwen3.6 27B
alibaba/qwen3.6-27b
Alibaba$0.60$3.60256K
Qwen3.6 35B A3B
alibaba/qwen3.6-35b-a3b
Alibaba$0.38$2.25256K
Qwen3.6 Flash
alibaba/qwen3.6-flash
Alibaba$0.25$1.501M
Qwen3.6 Max Preview
alibaba/qwen3.6-max-preview
Alibaba$1.30$7.80256K
Qwen3.6 Plus
alibaba/qwen3.6-plus
Alibaba$0.50$3.001M
Qwen3.7 Max Preview
alibaba/qwen3.7-max-preview
Alibaba$2.50$7.501M
DeepSeek V3.2
deepseek/deepseek-v3.2
DeepSeek$0.57$1.71
DeepSeek V4 Flash
deepseek/deepseek-v4-flash
DeepSeek$0.20$0.40
GLM-5.2 Fast Preview
zai-org/glm-5.2-fast-preview
Z.ai$2.80$8.80
IndicTrans2 EN→Indic 1B
ai4bharat/indictrans2-en-indic-1B
AI4Bharat$0.021$0.042
DistilBERT SST-2
huggingface/distilbert-sst-2-int8
Hugging Face
M2M100 1.2B
meta/m2m100-1.2b
Meta$0.021$0.042
By category

Shortcuts by what you're actually optimizing for.

Five opinionated slices: coding agents, reasoning-heavy work, cheapest model with tool calling, longest context, and fastest text models. Each slice cites the underlying benchmark or pricing field directly.

Best at coding · SWE-bench

ModelProviderSWE-bench
Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic72.5
Kimi-K2.6
moonshot/kimi-k2.6
Moonshot68.2
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic62.1

Best reasoning · GPQA / MMLU

ModelProviderScore
GPT-5.4
openai/gpt-5.4
OpenAI91.8 MMLU
Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic90.4 MMLU
Gemini 3.1 Pro
google/gemini-3.1-pro
Google89.6 MMLU
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic87.2 MMLU
GPT-5.4 Mini
openai/gpt-5.4-mini
OpenAI84.5 MMLU
Gemini 3 Flash
google/gemini-3-flash
Google82.3 MMLU
Claude Haiku 4.5
anthropic/claude-haiku-4.5
Anthropic80.1 MMLU

Cheapest with tool calling

ModelProviderIn + Out /M
Granite-4.0-H-Micro
ibm-granite/granite-4.0-h-micro
Ibm-granite$0.13
Qwen3-30b-A3b-Fp8
qwen/qwen3-30b-a3b-fp8
Alibaba Qwen$0.39
Gemma-4-26b-A4b-IT
google/gemma-4-26b-a4b-it
Google$0.40
GPT-5 Nano
openai/gpt-5-nano
OpenAI$0.45
Qwen Flash
alibaba/qwen-flash
Alibaba$0.45
Qwen3-VL Flash
alibaba/qwen3-vl-flash
Alibaba$0.45
Glm-4.7-Flash
zai-org/glm-4.7-flash
Zai-org$0.46
Gpt-Oss-20b
openai/gpt-oss-20b
OpenAI$0.50

Longest context window

ModelProviderContext
Grok 4.20 Multi-Agent
xai/grok-4.20-multi-agent-0309
xAI2M
Grok 4.20 Non-Reasoning
xai/grok-4.20-0309-non-reasoning
xAI2M
Grok 4.20 Reasoning
xai/grok-4.20-0309-reasoning
xAI2M
GPT-5.6 Luna
openai/gpt-5.6-luna
OpenAI1.1M
GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI1.1M
GPT-5.6 Terra
openai/gpt-5.6-terra
OpenAI1.1M
Gemini 3.5 Flash
google/gemini-3.5-flash
Google1.0M
Gemini 3.5 Flash-Lite
google/gemini-3.5-flash-lite
Google1.0M

Fastest · edge + mini tier

ModelProviderTier
Llama-3.3-70b-Instruct-Fp8-Fast
meta/llama-3.3-70b-instruct-fp8-fast
Metaedge
Llama-4-Scout-17b-16e-Instruct
meta/llama-4-scout-17b-16e-instruct
Metaedge
Grok 4 Fast
xai/grok-4-fast
xAIedge
Claude Haiku 4.5
anthropic/claude-haiku-4.5
Anthropicedge
GPT-4.1 Mini
openai/gpt-4.1-mini
OpenAIedge
GPT-4.1 Nano
openai/gpt-4.1-nano
OpenAIedge
GPT-4o Mini
openai/gpt-4o-mini
OpenAIedge
GPT-5 Mini
openai/gpt-5-mini
OpenAIedge
How this is built

Public benchmarks. Live pricing. Free JSON.

Benchmarks: sourced from each lab's published eval card (Anthropic system cards, OpenAI release notes, Google's technical reports, Moonshot's Kimi paper, Meta's Llama reports). Missing cells mean the lab hasn't disclosed that benchmark — not that the model failed it.

Pricing: the exact pass-through rate you pay on AIgateway. The $/M tokens here is what the provider charges — nothing added.

Updates: the catalog updates on every release, so this page moves when a new frontier model lands. The JSON endpoint is at GET /api/leaderboard with a one-hour browser cache — embed it on your own comparison post.

Want to run any of these? Every row links to a model page with quickstart code. Or try the cost calculator to project spend across candidate models for your workload.