OpenAI Reports Triple Performance Increase on ARC-AGI-3 Benchmark Through API Configuration

OpenAI achieved a threefold score improvement on the ARC-AGI-3 benchmark by enabling two specific API settings in GPT-5.6.

OpenAI has reported a significant performance improvement on the ARC-AGI-3 benchmark, achieving triple the baseline scores through two API configuration changes in GPT-5.6, according to the company’s announcement.

The performance gains were achieved by enabling two specific settings: retaining reasoning processes and enabling compaction. According to OpenAI, these settings not only boosted scores on the ARC-AGI-3 benchmark but also improved overall efficiency. The ARC-AGI-3 benchmark is designed to test abstract reasoning and general intelligence capabilities in AI systems.

The announcement did not provide detailed technical specifications about how the reasoning retention and compaction features function, but OpenAI indicated that the combination of these two settings was key to the dramatic improvement in performance metrics. The company emphasized that the changes improved both the accuracy of results and the efficiency of the model’s operation on this particular benchmark.