OpenAI triples ARC-AGI-3 score with "two settings": benchmark encounters crisis of trust

OpenAI released a research report on July 29, revealing that "two settings" can triple the model's score on the ARC-AGI-3 benchmark, revealing that the evaluation results are strongly related to the inference/decoding settings, and reminding the industry to be wary of falsely high results caused by "evaluation configuration dependence."

On July 29, OpenAI released a research report that seemed to "expose its own shortcomings": it revealed that through "two settings" the model's score on the ARC-AGI-3 benchmark could be increased to three times. If even "settings" can make a threefold difference, then how should we measure the real capability gap between models?

What is ARC-AGI-3

ARC-AGI (Abstraction and Reasoning Corpus for AGI) is an important benchmark for measuring a model's ability to "solve new problems/abstract reasoning" and is widely regarded as a benchmark for "human-like abstract reasoning." Its third generation (ARC-AGI-3) places higher requirements on the generalization and reasoning capabilities of the model. OpenAI found that certain inference/decoding settings (usually related to inference budget, search strategy, or sampling) had a huge impact on performance—adjusting just two items resulted in a threefold jump.

A methodological reminder

The core value of this report is not in the results, but in the methodological reminder: The "real capability" of the model is strongly related to the "evaluation setting", and the evaluation results must be under a unified configuration before they can be compared horizontally. This directly shakes the foundation of narratives such as "a certain model scores X" - if someone else uses a different inference budget to run a higher score, how convincing is this score?

The pertinence of this reminder is quite obvious. Just a few days ago (July 24), Anthropic, the old rival behind ChatGPT, claimed at the launch of Claude Opus 5 that its ARC-AGI 3 score was 3x higher than the next best model. The "offense and defense" between OpenAI and Anthropic around ARC-AGI-3 exactly reflects the status of the "Inference Baseline" in 2026 as the core of the model's ability narrative - but it also makes the question of "who scores high" more and more confusing.

From an industry perspective, this report reveals a spreading industry phenomenon: when benchmark scores become a marketing weapon, manufacturers will optimize "configuration" rather than "capabilities." This is an important warning for the domestic large model evaluation ecosystem - whether it is a public list or an internal evaluation, the "inference budget, sampling settings, and evaluation protocols" must be clear and unified, otherwise the scores will just become a meaningless numbers game. The true comparison of capabilities will eventually return to "reproducible evaluation under uniform conditions."

Several directions worth tracking in the future:

  1. The specific content of "two settings": Whether OpenAI will disclose the complete setting details for reproduction.
  2. Industry consensus on evaluation specifications: Whether there will be a unified ARC-AGI-3 evaluation protocol.
  3. Anthropic's response: How does Opus 5's ARC-AGI 3 score self-certify in the face of "configuration dependency" doubts.
  4. Revision of domestic evaluation standards: Whether the public list will tighten the disclosure requirements for evaluation configurations.
Copyright: Content sourced from OpenAI Official News (Research) . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...