← Latest reporting

A benchmark update should trigger a model-selection retest, not a leaderboard switch

Artificial Analysis changed the composition and weighting of its Intelligence Index. That is useful evidence, but enterprises should replay their own tasks before changing a model decision.

AI Capability FrontierSkills Systems and HR Tech
A flat screenprint shows three misaligned measuring frames crossed by the same test tile.
Conceptual AI illustration of a benchmark retest; it is not a factual chart.

What happened

Artificial Analysis published Intelligence Index v4.3.2 with 10 evaluations and a 30% weight for agent tasks.

Why it matters

A methodology change can move the comparison target even when the underlying model has not changed.

Artificial Analysis now describes Intelligence Index v4.3.2 as a weighted combination of 10 evaluations. Agent tasks carry 30% of the index, coding and scientific reasoning 20% each, and general tasks 30%. The publisher estimates the aggregate 95% confidence interval at under ±1%, while noting that individual evaluations may be wider and that the suite is primarily text-based and English-language.

That disclosure makes the index more useful, not universal. A new index version changes the measurement instrument: tasks, weights, judging methods and anchors can all shift. A model that rises after the change may fit the revised suite better without becoming better on a particular organisation's documents, languages, latency budget or failure costs.

Preserve the decision, then rerun it

Record the index version, model version, settings, price and evaluation date behind every selection decision. When the external methodology changes, do not overwrite the old result. Create a new decision record and replay a stable set of local tasks: successful cases, known failures, sensitive edge cases and representative production inputs. Compare quality, abstention, tool errors, latency and total cost.

A research paper on leaderboard sensitivity found that small changes to question order or answer selection could move rankings by as many as eight positions on common multiple-choice benchmarks. That result does not invalidate the new index: the current suite contains agentic, coding and open-answer tasks, and Artificial Analysis documents several controls. It does explain why a rank alone is a weak procurement instruction.

Set a change threshold

Before testing, define what would justify a switch. A one-point external movement should not automatically beat migration cost, new security review, altered data terms or a material regression in a critical task. Require a local improvement outside normal test variation and no new high-severity failure.

Use the Skills Intelligence methodology to keep external evidence, local evidence and the final decision separate. The correct response to a better benchmark is a better retest—not a reflexive vendor change.