OpenAI President Greg Brockman proudly declared “welcome to the AGI era” after the GPT-6 Astra release, but this claim stands in stark contrast to CEO Sam Altman’s earlier statement that AGI is an “irrelevant marketing term.”
The controversy centers on dramatically different benchmark results. GPT-6 Astra scored up to 99.9% on the ARC-AGI-3 test, but when all models were evaluated in a unified, neutral Standard Harness environment, the score plummeted to 62.7%.
In third-party organization Artificial Analysis’s “Intelligence Index 4.1.1,” GPT-6 Astra scored 61.2—only slightly above GPT-5.6 Sol’s 60.9, and below Claude Fable 5.1’s 65.7 and Claude Opus 5’s 63.1.
GPT-6 Astra achieved a significant leap in cybersecurity capabilities, scoring 100% on ExploitBench, becoming OpenAI’s first model to reach the “Critical” level in its Preparedness Framework for cybersecurity capabilities. OpenAI stated the model can autonomously find unknown security vulnerabilities and develop new exploitation methods without step-by-step human guidance.










Laat een reactie achter