LLM Leaderboard

Find the right LLM for your workload

Compare leading models across benchmark performance, speed, and cost to understand the trade-offs between them.

Best LLMs Per Task

See which models perform best across reasoning, coding, and tool use benchmarks.

Best LLMs Per Task

See which models perform best across reasoning, coding, and tool use benchmarks.

Reasoning (GPQA Diamond)

94.8%

Gemini 3.7 Flash (High)

94.60%

GPT-5.4 Pro (xhigh)

94.40%

Gemini 3.1 Pro Preview (high)

94.10%

Gemini 3.6 Flash (high)

94.00%

Grok 4.6 (high)

Coding (SWE-bench Verified)

83.50%

Claude Opus 4.7 (max)

80.60%

GPT-5.5 (xhigh)

79.30%

Gemini 3.5 Flash (high)

78.70%

Claude Opus 4.6 (no thinking)

78.70%

GLM-5.2 (max)

Tool Use (BFCLV4)

77.47%

Claude Opus 4.5 (FC)

73.24%

Claude Sonnet 4.5 (FC)

72.51%

Gemini 3 Pro Preview (Prompt)

72.38%

GLM-4.6 (FC thinking)

69.57%

Grok-4.1 Fast Reasoning (FC)

Fastest & Most Cost-Efficient LLMs

See how leading models differ in throughput, latency, and token pricing to understand the trade-offs between speed and cost.

Fastest & Most Cost-Efficient LLMs

See how leading models differ in throughput, latency, and token pricing to understand the trade-offs between speed and cost.

Fastest Models

1522

Celeris-1

909

Mercury 2

399

Step 3.7 Flash

370

Ling 3.0 Flash

368

Gemini 3.5 Flash-Lite

tokens/second

Lowest Latency

(TTFT)

0.4

Command A+

0.572

North Mini Code

0.6

Celeris-1

0.65

Ministral 3 3B

0.74

Gemini 3.7 Flash

Cheapest Models

Input Cost

Output Cost

$0.16

$0.12

$0.08

$0.04

$0

Gemma 4 E4B

Sarvam 30B

Nova Micro

Qwen3.5 4B

Nemotron Nano 9B V2

Model Comparison

Model Comparison

Claude Haiku 4.5Anthropic2025-10-15200,000$1.00$5.0037.40%73.30%68.7
Claude Opus 4.5Anthropic2025-11-24200,000$5.00$25.0087.00%77.47%N/A
Claude Opus 4.6Anthropic2026-02-051,000,000$5.00$25.0091.31%N/A78.7
Claude Opus 4.7Anthropic2026-04-161,000,000$5.00$25.0094.20%83.50%N/A
Claude Opus 5Anthropic2026-07-241,000,000$5.00$25.0093.90%73.30%N/A
Claude Sonnet 4.5Anthropic2025-09-291,000,000$3.00$15.0083.40%71.30%73.24%
GLM-4.6Z.ai2025-09-30200,000$0.50$2.2080.50%68.00%72.38%
GLM-5.2Z.ai2026-06-131,048,576$1.40$4.4091.20%78.70%N/A
GPT-5.2OpenAI2025-12-11400,000$1.75$14.0092.40%80.00%55.87%
GPT-5.4OpenAI2026-03-051,050,000$2.50$15.0092.80%77.20%N/A
GPT-5.4 ProOpenAI2026-03-051,050,000$30.00$180.0094.60%77.20%N/A
GPT-5.5OpenAI2026-04-231,050,000$5.00$30.0094%80.60%N/A
GPT-5.5 ProOpenAI2026-04-231,050,000$30.00$180.0093.90%N/AN/A
GPT-5.6 SolOpenAI2026-07-091,050,000$5.00$30.0094.10%N/AN/A
GPT-5.6 TerraOpenAI2026-07-091,050,000$2.00$12.0093.30%N/AN/A
Gemini 3 Pro PreviewGoogle2025-11-181,000,000$2.00$12.0092.50%76.20%72.51%
Gemini 3.1 Pro PreviewGoogle2026-02-191,000,000$2.00$12.0094.40%80.60%N/A
Gemini 3.5 FlashGoogle2026-05-191,000,000$1.50$9.0093.30%79.30%N/A
Gemini 3.6 FlashGoogle2026-07-211,048,576$1.50$3.7594.10%79.60%N/A
Gemini 3.7 FlashGoogle2026-08-131,000,000$0.75$3.7594.8%
Grok 4xAI2025-07-09256,000$2.00$6.0087.70%72.00%N/A
Grok 4.1 FastxAI2025-11-192,000,000$0.20$0.5085.30%60.00%69.57%
Grok 4.5xAI2026-07-08500,000$2.00$6.0093.40%N/AN/A
Grok 4.6xAI2026-08-12500,000$2.00$6.0094.00%N/AN/A
Kimi K2 InstructMoonshot AI2025-07-11256,000$0.50$2.0075.10%65.80%59.06%
Kimi K2.6Moonshot AI2026-04-20262,144$0.54$4.0090.50%80.20%N/A
Qwen3.6 Max PreviewAlibaba2026-04-20262,144$1.03$7.8088.80%72.80%N/A
Qwen3.7 MaxAlibaba2026-05-191,000,000$2.50$7.5092.40%80.40%N/A
o3OpenAI2025-04-16200,000$2.00$8.0087.70%71.70%63.05%

About benchmark scores

Benchmark results can vary depending on the model version, reasoning or thinking configuration, prompts, and other test settings. Where multiple configurations are available, we use the highest-scoring result reported by the selected benchmark source. Scores should therefore be used as comparative signals rather than exact measures of real-world performance.

About benchmark scores

Benchmark results can vary depending on the model version, reasoning or thinking configuration, prompts, and other test settings. Where multiple configurations are available, we use the highest-scoring result reported by the selected benchmark source. Scores should therefore be used as comparative signals rather than exact measures of real-world performance.

FAQ

Frequently asked questions

What is the LLM Leaderboard, and how does it work?

The LLM Leaderboard compares leading models across capability, speed, latency, context window, and cost. It brings benchmark and performance data into one place so you can compare models without relying on a single score.

What benchmarks are used to evaluate models on the LLM Leaderboard?

We use GPQA Diamond for reasoning, SWE-bench Verified for coding, and BFCL V4 for tool use. Each benchmark measures a different capability, so the scores are shown separately rather than combined into one overall ranking.

How does the LLM Leaderboard rank models based on speed, cost, and latency?

Speed is measured using median output throughput in tokens per second, while latency is measured by time to first token or first chunk. Cost comparisons use current input and output token pricing so you can see how model performance affects operating cost.

What details are included in the model comparison table?

The comparison table includes each model’s:

  • Provider

  • Release date

  • Context window

  • Input and output cost

  • Available benchmark scores (GPQA, SWE-bench Verified, BFCL V4)

Where a model hasn’t been evaluated by a benchmark source, the score is left blank rather than replaced with an incomparable result.

How can I use the LLM Leaderboard to choose the best model for my use case?

Start with the capability that matters most for your workload like reasoning, coding, or tool use. Then compare context limits, latency, throughput, and cost to find the model that gives you the best overall trade-off for your application.

Create an account and start building today.

Create an account and start building today.

Create an account and start building today.

Create an account and start building today.