OpenAI has announced that its latest artificial intelligence model, GPT-5.6 Sol, has outperformed Anthropic's rival Claude 5 Opus in the prestigious ARC-AGI-3 benchmark for artificial general intelligence. However, the impressive 38.3% score immediately sparked controversy as it was only achieved through OpenAI's proprietary custom evaluation system. In the official, standard testing environment, OpenAI's model recorded a highly modest score of just 7.8%.
Detailed Developments
The race for leadership in artificial general intelligence (AGI) between OpenAI and Anthropic continues to heat up with the latest announcement from the Sam Altman-led company. According to reports from The Decoder, OpenAI issued a direct response to Anthropic's previous record on the ARC-AGI-3 leaderboard. Anthropic had earlier boasted that its flagship model, Claude 5 Opus, achieved a score of 30.2% under standard testing conditions. To assert its dominance, OpenAI claimed that GPT-5.6 Sol reached 38.3%, officially overtaking its rival. Nevertheless, the massive gap between the custom environment results and the official ARC-AGI-3 environment has led industry observers to question the real-world significance of this achievement.
Technical Analysis & Technology
The core difference lies in the technical approach OpenAI applied to GPT-5.6 Sol during testing. Rather than running directly in the standard ARC-AGI-3 environment, OpenAI operated the model through a custom API integrated with specialized optimization technologies. Specifically, this system utilizes 'retained reasoning' and continuous 'context compaction' to prevent the model from losing track of complex data points during problem-solving tasks. These advanced software optimization techniques significantly boost LLM processing capabilities but are not permitted within the standard ARC-AGI evaluation framework. In contrast, Anthropic's Claude 5 Opus achieved its 30.2% benchmark purely on its baseline capabilities, without relying on any external assistance tools or special fine-tuning outside the official test environment.
Expert Opinions & Insights
Many industry experts note that OpenAI's claim reflects a concerning trend where AI research labs have begun 'over-optimizing' for benchmarks. Modifying the test environment to secure higher scores can distort the fundamental purpose of independent evaluations like ARC-AGI-3, which are designed to measure an AI's ability to learn new skills in novel situations. Specialists argue that without a standardized testing environment, comparative metrics between models will gradually lose their practical value for end-users and application developers.
Impact & Future Outlook
This development illustrates how AI benchmarks are turning into fierce marketing battlegrounds rather than objective technical measures. For the tech community both in Vietnam and globally, understanding the context behind benchmark figures is crucial to avoid falling for corporate PR hype. Moving forward, independent organizations behind benchmarks like ARC-AGI will likely need to tighten testing environment rules to ensure fairness. The battle for AGI supremacy remains far from over as competitors continue to deploy technical arguments to defend their claims.