Open vs Closed · benchmark standards

What the two AA index versions actually measure

The catch-up page reads every model on one of two fixed versions of the Artificial Analysis Intelligence Index: v4.1.1 (frozen on its last day, 2 Sep 2026) and v4.3.2 (captured 30 Sep 2026). They share a name and a 0–100 scale but not a test suite. This page lists what is inside each one, what changed between them, how the number is computed, and where every statement comes from. The main view uses v4.3.2, AA's current index. The other rulers on this page, including two we built from AA's own measurements (the fixed items and the legacy suite), exist only to check it.

At a glance

AA's two versions, side by side

AgentsCoding Scientific reasoningGeneral
Weights

Four rulers, one table

Every ruler is the same kind of number: 100 × Σ (weight × item score), each item on a 0–1 scale, the weights in a column summing to 100%. The rulers differ only in which items they take, which version of each item, and the weights. A blank cell means the item is not in that ruler.

What the version change does

AA's September suite gives OpenAI and Anthropic extra points

Cross-check only

Legacy suite (measured only, to May 2026)

Which models it covers

Item by item

What carried over, what changed

"Identical" is not taken from AA's changelog: it is checked on the data. For every configuration that has the item on both 2 Sep (v4.1.1) and 30 Sep (v4.3.2), the two values are compared digit for digit. An item that AA says is unchanged but whose values moved would show up here.

Full definition

v4.1.1: the ten items

v4.3.2: the eleven items

Arithmetic

How the 0–100 number is built

Elo-type evaluations, in AA's words

Can the published number be rebuilt from the published items?

The weak spot

Estimated scores

Our own scale

Fixed items (cross-check, all measured, up to today)

History

Every index version since 2024

Sources

Where each statement comes from