Cortex Lab

LLM performance leaderboard

Price, output speed and quality across every major model, compared side by side. This is the table we open before deciding what a workflow should run on. The data is published by Artificial Analysis — we host it here for convenience and credit it in full.

120 models
Quality index Cost $/1M tokens Speed Benchmarks
# Model
1 Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic Sep 2026 57.6 — — $8.00 $4.00 $20.00 96 477.23 477.23 — 61.4% 66.9% 84.7% — — — — — — — —
2 Claude Opus 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) Anthropic Sep 2026 56.0 — — $8.00 $4.00 $20.00 81 53.97 53.97 — 57.5% 65.0% 84.7% — — — — — — — —
3 Claude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic Sep 2026 56.0 — — $4.00 $2.00 $10.00 145 308.29 308.29 — 55.0% 61.0% 82.7% — — — — — — — —
4 Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback) Anthropic Sep 2026 53.6 — — $8.00 $4.00 $20.00 74 15.65 15.65 — 55.6% 60.4% 82.7% — — — — — — — —
5 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) Anthropic Sep 2026 53.4 81.6 — $20.00 $10.00 $50.00 69 114.39 114.39 93.7% 59.1% 63.1% 85.3% — — — — — — 91.4% 47.2%
6 Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) Anthropic Sep 2026 53.2 80.7 — $20.00 $10.00 $50.00 60 34.30 34.30 93.4% 58.7% 60.9% 83.0% — — — — — — 91.0% 45.8%
7 GPT-6 Astra (max) OpenAI Sep 2026 52.7 76.9 — $20.00 $10.00 $50.00 59 177.90 177.90 96.1% 54.7% 56.5% 80.7% — — — — — — 88.4% 41.4%
8 GPT-6 Astra (xhigh) OpenAI Sep 2026 52.4 75.9 — $20.00 $10.00 $50.00 54 74.01 74.01 96.3% 54.6% 55.7% 80.0% — — — — — — 89.1% 43.1%
9 Claude Sonnet 5.5 (Adaptive Reasoning, Xhigh Effort, Default Fallback) Anthropic Sep 2026 51.9 — — $4.00 $2.00 $10.00 110 15.04 15.04 — 50.0% 57.3% 79.7% — — — — — — — —
10 GPT-6.1 Sol (max) OpenAI Sep 2026 51.8 — — $4.00 $2.00 $10.00 87 118.51 118.51 — 52.9% 54.2% 83.0% — — — — — — — —
11 Claude Opus 5.5 (Adaptive Reasoning, Medium Effort, Default Fallback) Anthropic Sep 2026 51.2 — — $8.00 $4.00 $20.00 74 12.75 12.75 — 54.7% 59.3% 84.3% — — — — — — — —
12 Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) Anthropic Sep 2026 51.2 79.1 — $20.00 $10.00 $50.00 51 6.23 6.23 90.6% 55.9% 58.7% 83.7% — — — — — — 89.9% 43.1%
13 GPT-6.1 Sol (xhigh) OpenAI Sep 2026 51.0 — — $4.00 $2.00 $10.00 89 37.54 37.54 — 52.6% 55.7% 79.7% — — — — — — — —
14 GPT-6 Astra (high) OpenAI Sep 2026 50.9 77.1 — $20.00 $10.00 $50.00 52 29.69 29.69 94.9% 53.1% 55.4% 80.0% — — — — — — 89.9% 40.0%
15 Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic Jul 2026 50.8 78.0 — $10.00 $5.00 $25.00 — — — 93.2% 54.9% 56.4% 79.3% — — — — — — 89.1% 42.1%
16 GPT-6.1 Sol (high) OpenAI Sep 2026 50.2 — — $4.00 $2.00 $10.00 78 19.18 19.18 — 51.4% 55.8% 82.3% — — — — — — — —
17 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Anthropic Jul 2026 49.7 77.0 — $10.00 $5.00 $25.00 — — — 93.7% 54.4% 55.7% 80.3% — — — — — — 88.0% 43.3%
18 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic Jun 2026 49.6 76.5 — $20.00 $10.00 $50.00 — — — 92.6% 55.5% 61.0% 82.3% 63.5% 98.5% 62.9% — — — 84.6% 38.1%
19 GPT-6 Astra (medium) OpenAI Sep 2026 49.6 76.7 — $20.00 $10.00 $50.00 52 3.58 3.58 93.9% 52.7% 54.2% 79.7% — — — — — — 89.5% 35.5%
20 Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) Anthropic Sep 2026 48.9 77.1 — $20.00 $10.00 $50.00 50 4.80 4.80 88.6% 53.8% 56.4% 84.7% — — — — — — 88.0% 41.0%
21 Claude Opus 5 (Adaptive Reasoning, High Effort) Anthropic Jul 2026 48.1 76.5 — $10.00 $5.00 $25.00 — — — 93.7% 52.8% 55.4% 79.0% — — — — — — 87.6% 44.7%
22 Muse Spark 1.3 (max) Meta Sep 2026 48.1 75.8 — $2.00 $1.25 $4.25 226 20.92 29.78 93.5% 48.7% 58.8% 83.0% — — — — — — 84.3% 50.5%
23 GPT-6.1 Sol (medium) OpenAI Sep 2026 47.8 — — $4.00 $2.00 $10.00 74 4.13 4.13 — 49.9% 53.2% 83.3% — — — — — — — —
24 GPT-6 Sol (max) OpenAI Sep 2026 47.5 — — $4.00 $2.00 $10.00 83 108.52 108.52 — 47.9% 57.6% 83.7% — — — — — — — —
25 GPT-5.6 Sol (max) OpenAI Jul 2026 47.0 77.4 — $8.00 $4.00 $20.00 — — — 94.1% 49.5% 57.1% 84.0% 72.7% 85.1% 65.9% — — — 88.0% 44.3%
26 Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) Anthropic Sep 2026 46.8 75.2 — $20.00 $10.00 $50.00 51 2.95 2.95 88.1% 48.9% 56.7% 82.3% — — — — — — 85.0% 39.0%
27 Claude Sonnet 5.5 (Adaptive Reasoning, High Effort, Default Fallback) Anthropic Sep 2026 46.7 — — $4.00 $2.00 $10.00 106 7.65 7.65 — 45.8% 53.7% 78.0% — — — — — — — —
28 Grok 4.7 (xhigh) SpaceXAI Sep 2026 46.4 — — $3.00 $2.00 $6.00 83 19.84 19.84 — 43.1% 57.4% 76.7% — — — — — — — —
29 Grok 4.7 (high) SpaceXAI Sep 2026 46.3 — — $3.00 $2.00 $6.00 79 20.37 20.37 — 42.3% 57.8% 77.0% — — — — — — — —
30 MiMo-V2.6-Pro Xiaomi Sep 2026 46.3 — — $0.54 $0.44 $0.87 43 3.11 49.73 — 49.4% 60.9% 86.3% — — — — — — — —

Independent benchmarks & pricing of LLMs — we publish them here, we don’t produce them. For further analysis and methodology, see artificialanalysis.ai. Updated 30 Sep 2026.

How to read it

Three things that trip people up when they first use this table.

Price is per million tokens

Input and output are priced separately, and output is usually the expensive half. A workflow that reads a lot and writes a little costs far less than the headline number suggests.

Two different kinds of speed

Time to first token is what a person waiting on a chat reply feels. Tokens per second is what matters for a batch job running overnight. Optimising for the wrong one wastes money.

Quality scores are a shortlist, not a verdict

Aggregate indices are a good way to narrow twenty models to three. They are a poor way to choose between those three — for that, test them on your actual task.

Building a workflow? We pick the model for you

We cost and benchmark every automation before it ships — so it runs on the cheapest model that actually holds up on your task.