The board.
28 models omni can actually run, each scored on 5 benchmarks against the dollars and the seconds a task was measured to take. Pick a board and the left edge of it becomes the dial: a model earns a rung when nothing else is both better and cheaper.
Turn a provider off. The edge moves, and models another vendor was shadowing step straight back onto it — which is why the same number means something different to somebody signed into fewer things than you.
Change the benchmark and the board itself changes shape. A model that owns one is mid-table on the next. That is the whole reason the registry asks what you are doing before it answers, instead of keeping one league table and calling it intelligence.
Six words, and who answers them.
Chosen by hand, best first. The first vendor on a list you are signed into answers, so losing one walks you a place down it. We tried working this out arithmetically and it produced answers that were defensible and wrong.
| ask for | when | in order |
|---|---|---|
fast | answer soonest. the quickest thing each vendor has, running hot | gpt-6-astra low fast › gemini-3.8-flash-low › claude-opus-5 low fast |
code | write and change code | gpt-6-astra high › claude-opus-5 max › gemini-3.8-flash-high |
design | make something somebody has to look at | claude-opus-5 max › gpt-6-astra high › gemini-3.8-flash-high |
research | read a lot, and be right | gpt-6-astra high › claude-opus-5 max › gemini-3.8-flash-high |
cost | spend as little as the job allows | gpt-6-astra low › gemini-3.8-flash-low › claude-opus-5 low |
general | no strong opinion. a good default | gpt-6-astra high › claude-opus-5 medium › gemini-3.8-flash-high |
Five leaderboards, and what each of them has actually measured.
A benchmark only earns a place here if somebody published the number. Coverage is uneven and stays uneven: a model no board has scored carries no score, and is simply not a candidate on that board.
| ask for | leaderboard | scored |
|---|---|---|
artificial-analysis-intelligence-index | Artificial Analysis Intelligence Index v4.3 | 28 of 28 |
gdpval-aa | GDPval-AA v2 | 28 of 28 |
terminal-bench | Terminal-Bench 4.0 | 26 of 28 |
agents-last-exam | Agents' Last Exam | 3 of 28 |
design-arena-full-stack | Design Arena — Full Stack | 18 of 28 |
The database, as it stands this second.
One row per model and effort, because effort changes all three numbers. A dash is a board that has not scored that model, and it is left as a dash rather than estimated onto the chart.
| model | artificial-analysis-intelligence-index | gdpval-aa | terminal-bench | agents-last-exam | design-arena-full-stack | $/task | time |
|---|---|---|---|---|---|---|---|
| gpt-6-astra max | 52.8 | 1580 | 0.59 | — | — | $3.26 | 8m 12s |
| gpt-6-astra xhigh | 52.5 | 1555 | 0.6 | — | — | $2.31 | 5m 38s |
| gpt-6-astra high | 51 | 1530 | 0.54 | — | — | $1.72 | 4m 05s |
| claude-opus-5 max | 50.7 | 1735 | 0.49 | 49 | 6.2 | $5.86 | 14m 12s |
| gpt-6-astra medium | 49.7 | 1501 | 0.49 | — | — | $1.54 | 3m 41s |
| claude-opus-5 xhigh | 49.7 | 1708 | 0.46 | — | 6.8 | $4.88 | 12m 22s |
| claude-opus-5 high | 48.2 | 1629 | 0.46 | — | 6.2 | $3.61 | 9m 19s |
| gpt-6-astra low | 46 | 1419 | 0.42 | — | — | $0.8175 | 1m 36s |
| claude-opus-5 medium | 45.1 | 1525 | 0.34 | — | 6.2 | $2.19 | 5m 51s |
| gpt-5.6-terra max | 42.3 | 1477 | 0.35 | 50.7 | 6.7 | $1.40 | 6m 19s |
| gemini-3.8-flash-high | 41.2 | 1464 | 0.2 | — | — | $1.24 | 4m 01s |
| gemini-3.8-flash-medium | 40 | 1455 | 0.2 | — | — | $0.9310 | 0s |
| claude-opus-5 low | 39.8 | 1372 | 0.26 | — | 6.2 | $1.10 | 2m 54s |
| gpt-5.4 xhigh | 39 | 1307 | — | — | 3 | $0.0000 | 0s |
| gpt-5.5 xhigh | 38.6 | 1396 | 0.15 | 47.9 | 7 | $2.63 | 4m 28s |
| claude-sonnet-5 max | 38.4 | 1501 | 0.14 | — | 5.6 | $5.09 | 15m 39s |
| gpt-5.6-terra xhigh | 38.2 | 1479 | 0.1 | — | 6.7 | $0.6318 | 3m 37s |
| gpt-5.5 high | 37.3 | 1372 | 0.09 | — | 7 | $1.54 | 3m 17s |
| gpt-5.6-terra high | 34.5 | 1415 | 0.02 | — | 6.7 | $0.3379 | 2m 09s |
| gemini-3.6-flash-high | 34.3 | 1331 | 0.07 | — | 6.7 | $0.9288 | 3m 35s |
| gpt-5.5 medium | 34.2 | 1288 | 0.05 | — | 7 | $0.9017 | 1m 42s |
| gemini-3.8-flash-low | 33.8 | 1359 | 0.1 | — | — | $0.0000 | 0s |
| gemini-3.5-flash-high | 33 | 1259 | 0.07 | — | 6 | $1.56 | 4m 06s |
| gpt-5.6-terra medium | 32.8 | 1319 | 0.01 | — | 6.7 | $0.0000 | 0s |
| gpt-5.5 low | 30.7 | 1113 | — | — | 7 | $0.0000 | 0s |
| gpt-5.6-terra low | 27.9 | 1179 | 0.02 | — | 6.7 | $0.0000 | 0s |
| gpt-5.4-mini xhigh | 24.6 | 1095 | 0.02 | — | — | $0.4097 | 4m 44s |
| claude-haiku-4-5-20251001 | 17.6 | 854 | 0 | — | — | $0.2077 | 2m 48s |
The whole database is one file.
Editing it and committing is the whole publishing flow. No database, no login, no release of the library, no redeploy. Git does what a table would have done: history, blame, review, rollback — and it also means somebody's commit is your routing table.
Edit models.json and commit. This page reads it at request time, so the change is live on the next request and the library picks it up within the hour.
- Artificial Analysis measures first-party API configurations, which may differ from the CLI harnesses used here.
- Price and time use Intelligence Index v4.3 tasks. Missing measurements remain null and are excluded from numerical scales, but explicit keyword picks remain runnable.
- Astra ultra is supported by the app server but has no separately measured row here; its score is not inferred from max.
- Keyword shortlists are curated routing preferences. They preserve the requested Astra low/high and Flash 3.8 choices; they are not numerical benchmark rankings.
- Supplemental design and research measurements are being audited separately; no previous-model scores are transferred to Astra or Flash 3.8.