Practical intelligence index / 2026

The benchmark for models that do the whole job.

A clear view of model performance across coding, reasoning and agentic work — with cost and time kept in the frame.

Best composite score
78.4%
GPT-5.1 Codex-Max · agentic coding
vs. previous index+4.8 pts
Lowest cost / task
$0.09
DeepSeek-V4-Pro · median run
efficiency signaltop 12%
Runs completed
1,184
Across 9 task families

Leaderboard

Model configurations ranked by the current composite index

How scores work
#ModelProviderFocusScoreCostTrend
01
GPT-5.1 Codex-Maxhigh effort · tool use
OpenAIcoding78.4%$4.20+4.8
02
Claude Opus 4.1max effort · agent
Anthropicagentic74.1%$11.84+2.1
03
Gemini 3.7 Flashhigh thinking · tools
Googlereasoning71.8%$2.18+3.5
04
DeepSeek-V4-Prostandard · coding
DeepSeekcoding68.9%$0.09+1.7
05
Qwen3-Coder-480B-A35B-Instructmax context · coding
Qwencoding65.4%$0.42+1.2
06
Grok 4.6high reasoning · tools
xAIagentic63.2%$2.00+0.8
Showing 6 of 42 configurations

What makes a useful benchmark?

A score should reveal how a model behaves when the obvious answer is not the finished work.

01

Outcome over syntax

We score observable behavior and verified results, not similarity to a reference patch.

02

Context is part of the task

Repositories, documents, tools and changing constraints are part of the environment.

03

Quality has a price tag

Every run keeps time, tokens and cost visible beside the pass rate.

Run history

Recent suite activity across the indexed configurations

Latest batch
#0148
113 tasks · completed 12m ago
verified passes76.2%
Median latency
17min
Across all agent runs
vs. last week−8.4%
Next refresh
04:12h
Scheduled index update
data pipelinehealthy
Comparison queued