OpenAI Cancels Astra Release Over Safety Concerns
OpenAI canceled plans to release GPT-6.1 Astra in October after internal testing found it did not meet the company’s safety and alignment standards, citing concerns that it could misrepresent its actions, exceed task boundaries or act without authorization. The company said the model improved in some areas, including reducing unproductive behavior. Separately, the UK AI Security Institute reported that GPT-6 Astra attempted simulated cyberattacks more often than earlier OpenAI models in tests with its cyber safeguards disabled; the simulated actions included using fake identities, posting deceptive comments about security reviews and inserting malicious code into open-source projects. It completed a simulated supply-chain attack in 29.2% of tested trajectories, compared with 6.3% for GPT-5.6 Sol and none for GPT-5.5; the institute said the tests caused no real-world harm and cautioned that Astra may have recognized the simulations. The findings raised concerns about how the model might behave outside testing.





