Open vs Closed · data note

When do open weights catch the closed labs?

Twenty months of the Artificial Analysis leaderboard, rebuilt from archived snapshots plus today's live board, with every model placed on its release date and scored on one fixed ruler at a time. Part one is the race at the top and how long open weights take to reach each closed record. Part two is the price–intelligence frontier: at each budget, who sells the best model.

Ruler
AA estimates
Measurement error

Part 1 · The gap at the top

How far behind is the best open model?

Each line is the highest score available on that day, open weights (green) against closed (purple). A closed model counts from its release date; an open model counts from the day its weights were open, which is sometimes later. Pick any day, including days that have not happened yet. Past the dashed line the numbers come from the fitted curves, not from data.

Jump to
Best closed Best open (≥16B) Fitted curve (recency-weighted exponential) Gap Score is AA's estimate
Fit memory (half-life)

What the fit says, and how much to trust it

All possible crossing points (3 fits × 3 half-lives × 4 rulers, with and without the newest records), and why the fit is fragile

    How long until open reaches each closed record?

    Closed record released First open model at or above it Not reached yet (still counting) Estimated score

    Each lab's flagship, one generation at a time

    The chart above follows the global record, which one lab can hold for months. This table fixes a target per lab and generation: the best configuration of that flagship on its launch day, scored on the selected ruler. The clock stops at the first open-weight model (≥16B) at or above it. 0 d means an open model was already there on launch day.

    Every score in these tables, the test items behind it and the archived page it comes from: trace list (the trace link on each row opens that row). What the rulers contain: benchmark standards.

    Part 2 · The Pareto frontier

    At each price, who sells the best model?

    Every dot is a model on sale that week: its list price (3:1 input/output blend, per million tokens, log scale) against its score on the selected ruler. A model is on the frontier when nothing cheaper scores as high. Everything in the grey region is beaten by something that is both cheaper and at least as smart.

    Open weights Closed On the frontier Score is AA's estimate Dominated

    Share of frontier seats held by open weights

    New on the board

    What arrived since 2 September

    Show all
    Method

    Why four rulers, and not one long line

    Item-by-item definition of both AA versions, what changed between them and how the number is computed: benchmark standards.

    Sources