Twenty months of the Artificial Analysis leaderboard, rebuilt from archived snapshots plus today's live board, with every model placed on its release date and scored on one fixed ruler at a time. Part one is the race at the top and how long open weights take to reach each closed record. Part two is the price–intelligence frontier: at each budget, who sells the best model.
Each line is the highest score available on that day, open weights (green) against closed (purple). A closed model counts from its release date; an open model counts from the day its weights were open, which is sometimes later. Pick any day, including days that have not happened yet. Past the dashed line the numbers come from the fitted curves, not from data.
The chart above follows the global record, which one lab can hold for months. This table fixes a target per lab and generation: the best configuration of that flagship on its launch day, scored on the selected ruler. The clock stops at the first open-weight model (≥16B) at or above it. 0 d means an open model was already there on launch day.
Every score in these tables, the test items behind it and the archived page it comes from: trace list (the trace link on each row opens that row). What the rulers contain: benchmark standards.
Every dot is a model on sale that week: its list price (3:1 input/output blend, per million tokens, log scale) against its score on the selected ruler. A model is on the frontier when nothing cheaper scores as high. Everything in the grey region is beaten by something that is both cheaper and at least as smart.
Item-by-item definition of both AA versions, what changed between them and how the number is computed: benchmark standards.