On September 3, 2026, OpenAI released the results of GPT-6 model in the Astra benchmark test, and subsequently made multiple adjustments to the relevant data. The hallucination rate of the model decreased from 4.2% at first to 2%, then rose back to 4.2%. Metrics such as mathematics, cybersecurity, and code capabilities also underwent repeated adjustments. OpenAI stated that the adjustments were intended to ensure that the numbers represented the best estimate of the model’s available performance, and noted that test conditions (such as tool frameworks) could affect the results. This incident caused fluctuations in some results of Fable 5.1 and Opus 5 under competitor Anthropic. This situation sparked discussions in the industry, with some experts believing that repeated adjustments to test conditions might be for marketing purposes, and that insufficient technical details were disclosed. Previously, Meta was also accused of manipulating the evaluation results of Llama 4, and the industry is facing a long-term challenge of how to standardize measures of large model capabilities.
OpenAI revised the GPT-6 Astra benchmark data multiple times, with some scores changing significantly. Since the first announcement on the afternoon of September 3 local time, Astra’s hallucination rate dropped from 4.2% to 2% and then rose back to 4.2%. Its mathematical capabilities, cybersecurity, and code-solving abilities also underwent several adjustments. OpenAI stated that the revisions were made to ensure that the data accurately represents the model’s available performance, allowing users to make meaningful comparisons. They emphasized that test conditions, such as tool frameworks, can affect the results. Meanwhile, some scores of OpenAI’s competitor Anthropic models, Fable 5.1 and Opus 5, also fluctuated. This incident sparked discussions in the industry about the phenomenon of “scoring manipulation”; some experts believe that repeated adjustments to test conditions may be for marketing purposes, and there is insufficient disclosure of related technical details. Previously, Meta was also accused of modifying the evaluation scores of Llama 4, and the industry is facing a long-term challenge of how to standardize measures of large model capabilities.