Practical intelligence index / 2026
The benchmark for models that do the whole job.
A clear view of model performance across coding, reasoning and agentic work — with cost and time kept in the frame.
Best composite score
78.4%
GPT-5.1 Codex-Max · agentic coding
Lowest cost / task
$0.09
DeepSeek-V4-Pro · median run
Runs completed
1,184
Across 9 task families
Leaderboard
Model configurations ranked by the current composite index
How scores work →
#ModelProviderFocusScoreCostTrend
01OAIGPT-5.1 Codex-Maxhigh effort · tool use
OpenAIcoding78.4%$4.20+4.8
02ANTClaude Opus 4.1max effort · agent
Anthropicagentic74.1%$11.84+2.1
03GGemini 3.7 Flashhigh thinking · tools
Googlereasoning71.8%$2.18+3.5
04DSDeepSeek-V4-Prostandard · coding
DeepSeekcoding68.9%$0.09+1.7
05QWQwen3-Coder-480B-A35B-Instructmax context · coding
Qwencoding65.4%$0.42+1.2
06xAIGrok 4.6high reasoning · tools
xAIagentic63.2%$2.00+0.8
What makes a useful benchmark?
A score should reveal how a model behaves when the obvious answer is not the finished work.
01Outcome over syntax
We score observable behavior and verified results, not similarity to a reference patch.
02Context is part of the task
Repositories, documents, tools and changing constraints are part of the environment.
03Quality has a price tag
Every run keeps time, tokens and cost visible beside the pass rate.
Run history
Recent suite activity across the indexed configurations
Latest batch
#0148
113 tasks · completed 12m ago
Median latency
17min
Across all agent runs
Next refresh
04:12h
Scheduled index update