OpenAI triples ARC-AGI-3 score with "two settings": benchmark encounters crisis of trust
OpenAI released a research report on July 29, revealing that "two settings" can triple the model's score on the ARC-AGI-3 benchmark, revealing that the evaluation results are strongly related to the inference/decoding settings, and reminding the industry to be wary of falsely high results caused by "evaluation configuration dependence."
On July 29, OpenAI released a research report that seemed to "expose its own shortcomings": it revealed that through "two settings" the model's score on the ARC-AGI-3 benchmark could be increased to three times. If even "settings" can make a threefold difference, then how should we measure the real capability gap between models?
What is ARC-AGI-3
ARC-AGI (Abstraction and Reasoning Corpus for AGI) is an important benchmark for measuring a model's ability to "solve new problems/abstract reasoning" and is widely regarded as a benchmark for "human-like abstract reasoning." Its third generation (ARC-AGI-3) places higher requirements on the generalization and reasoning capabilities of the model. OpenAI found that certain inference/decoding settings (usually related to inference budget, search strategy, or sampling) had a huge impact on performance—adjusting just two items resulted in a threefold jump.
A methodological reminder
The core value of this report is not in the results, but in the methodological reminder: The "real capability" of the model is strongly related to the "evaluation setting", and the evaluation results must be under a unified configuration before they can be compared horizontally. This directly shakes the foundation of narratives such as "a certain model scores X" - if someone else uses a different inference budget to run a higher score, how convincing is this score?
The pertinence of this reminder is quite obvious. Just a few days ago (July 24), Anthropic, the old rival behind ChatGPT, claimed at the launch of Claude Opus 5 that its ARC-AGI 3 score was 3x higher than the next best model. The "offense and defense" between OpenAI and Anthropic around ARC-AGI-3 exactly reflects the status of the "Inference Baseline" in 2026 as the core of the model's ability narrative - but it also makes the question of "who scores high" more and more confusing.
From an industry perspective, this report reveals a spreading industry phenomenon: when benchmark scores become a marketing weapon, manufacturers will optimize "configuration" rather than "capabilities." This is an important warning for the domestic large model evaluation ecosystem - whether it is a public list or an internal evaluation, the "inference budget, sampling settings, and evaluation protocols" must be clear and unified, otherwise the scores will just become a meaningless numbers game. The true comparison of capabilities will eventually return to "reproducible evaluation under uniform conditions."
Several directions worth tracking in the future:
- The specific content of "two settings": Whether OpenAI will disclose the complete setting details for reproduction.
- Industry consensus on evaluation specifications: Whether there will be a unified ARC-AGI-3 evaluation protocol.
- Anthropic's response: How does Opus 5's ARC-AGI 3 score self-certify in the face of "configuration dependency" doubts.
- Revision of domestic evaluation standards: Whether the public list will tighten the disclosure requirements for evaluation configurations.
Reviews