silicon omni

The board.

28 models omni can actually run, each scored on 5 benchmarks against the dollars and the seconds a task was measured to take. Pick a board and the left edge of it becomes the dial: a model earns a rung when nothing else is both better and cheaper.

22 scored on GDPval-AA v2
$191312951677gdpval-aa ↑76410398210
claude-opus-5 max1735 on GDPval-AA v2 · $5.86 · 14m 12slevel 10 of 10 rungs · 22 plotted

Turn a provider off. The edge moves, and models another vendor was shadowing step straight back onto it — which is why the same number means something different to somebody signed into fewer things than you.

Change the benchmark and the board itself changes shape. A model that owns one is mid-table on the next. That is the whole reason the registry asks what you are doing before it answers, instead of keeping one league table and calling it intelligence.

01the shortlistsof 4

Six words, and who answers them.

Chosen by hand, best first. The first vendor on a list you are signed into answers, so losing one walks you a place down it. We tried working this out arithmetically and it produced answers that were defensible and wrong.

ask forwhenin order
fastanswer soonest. the quickest thing each vendor has, running hotgpt-6-astra low fast › gemini-3.8-flash-low › claude-opus-5 low fast
codewrite and change codegpt-6-astra high › claude-opus-5 max › gemini-3.8-flash-high
designmake something somebody has to look atclaude-opus-5 max › gpt-6-astra high › gemini-3.8-flash-high
researchread a lot, and be rightgpt-6-astra high › claude-opus-5 max › gemini-3.8-flash-high
costspend as little as the job allowsgpt-6-astra low › gemini-3.8-flash-low › claude-opus-5 low
generalno strong opinion. a good defaultgpt-6-astra high › claude-opus-5 medium › gemini-3.8-flash-high
02the boardsof 4

Five leaderboards, and what each of them has actually measured.

A benchmark only earns a place here if somebody published the number. Coverage is uneven and stays uneven: a model no board has scored carries no score, and is simply not a candidate on that board.

ask forleaderboardscored
artificial-analysis-intelligence-indexArtificial Analysis Intelligence Index v4.328 of 28
gdpval-aaGDPval-AA v228 of 28
terminal-benchTerminal-Bench 4.026 of 28
agents-last-examAgents' Last Exam3 of 28
design-arena-full-stackDesign Arena — Full Stack18 of 28
03every rowof 4

The database, as it stands this second.

One row per model and effort, because effort changes all three numbers. A dash is a board that has not scored that model, and it is left as a dash rather than estimated onto the chart.

modelartificial-analysis-intelligence-indexgdpval-aaterminal-benchagents-last-examdesign-arena-full-stack$/tasktime
gpt-6-astra max52.815800.59——$3.268m 12s
gpt-6-astra xhigh52.515550.6——$2.315m 38s
gpt-6-astra high5115300.54——$1.724m 05s
claude-opus-5 max50.717350.49496.2$5.8614m 12s
gpt-6-astra medium49.715010.49——$1.543m 41s
claude-opus-5 xhigh49.717080.46—6.8$4.8812m 22s
claude-opus-5 high48.216290.46—6.2$3.619m 19s
gpt-6-astra low4614190.42——$0.81751m 36s
claude-opus-5 medium45.115250.34—6.2$2.195m 51s
gpt-5.6-terra max42.314770.3550.76.7$1.406m 19s
gemini-3.8-flash-high 41.214640.2——$1.244m 01s
gemini-3.8-flash-medium 4014550.2——$0.93100s
claude-opus-5 low39.813720.26—6.2$1.102m 54s
gpt-5.4 xhigh391307——3$0.00000s
gpt-5.5 xhigh38.613960.1547.97$2.634m 28s
claude-sonnet-5 max38.415010.14—5.6$5.0915m 39s
gpt-5.6-terra xhigh38.214790.1—6.7$0.63183m 37s
gpt-5.5 high37.313720.09—7$1.543m 17s
gpt-5.6-terra high34.514150.02—6.7$0.33792m 09s
gemini-3.6-flash-high 34.313310.07—6.7$0.92883m 35s
gpt-5.5 medium34.212880.05—7$0.90171m 42s
gemini-3.8-flash-low 33.813590.1——$0.00000s
gemini-3.5-flash-high 3312590.07—6$1.564m 06s
gpt-5.6-terra medium32.813190.01—6.7$0.00000s
gpt-5.5 low30.71113——7$0.00000s
gpt-5.6-terra low27.911790.02—6.7$0.00000s
gpt-5.4-mini xhigh24.610950.02——$0.40974m 44s
claude-haiku-4-5-20251001 17.68540——$0.20772m 48s
04the sourceof 4

The whole database is one file.

Editing it and committing is the whole publishing flow. No database, no login, no release of the library, no redeploy. Git does what a table would have done: history, blame, review, rollback — and it also means somebody's commit is your routing table.

Edit models.json and commit. This page reads it at request time, so the change is live on the next request and the library picks it up within the hour.

  • Artificial Analysis measures first-party API configurations, which may differ from the CLI harnesses used here.
  • Price and time use Intelligence Index v4.3 tasks. Missing measurements remain null and are excluded from numerical scales, but explicit keyword picks remain runnable.
  • Astra ultra is supported by the app server but has no separately measured row here; its score is not inferred from max.
  • Keyword shortlists are curated routing preferences. They preserve the requested Astra low/high and Flash 3.8 choices; they are not numerical benchmark rankings.
  • Supplemental design and research measurements are being audited separately; no previous-model scores are transferred to Astra or Flash 3.8.