OpenAI faces accusations of violating California AI safety law with latest model releases

Logo via Wikimedia Commons; treatment-A cover, license to verify on approval

OpenAI faces accusations of violating California AI safety law with latest model releases

The Midas Project says OpenAI failed to publish required risk tiers for GPT 5.6 and Astra despite commitments in its California safety framework.

OpenAI has been accused of repeatedly violating California’s frontier AI safety law by failing to publish required risk assessments for several major model releases, including its latest Astra model.

The allegations come from the Midas Project, an AI industry watchdog that says OpenAI has failed to assign the risk tiers required under its own Frontier Governance Framework for at least three model releases this year.

California’s Transparency in Frontier AI Act, known as SB 53 before becoming law, requires major AI developers to publish safety frameworks explaining how they assess and mitigate advanced AI risks. Companies are then required to follow the policies they set out in those frameworks.

OpenAI published its Frontier Governance Framework in May. The framework says new models will be evaluated across four risk categories: cyber offense, chemical, biological, radiological and nuclear risks, harmful manipulation and loss of control. Models are supposed to receive a risk tier from one to three in each category.

Advertisement

The Midas Project says OpenAI has not publicly assigned those tiers to GPT 5.6 preview, GPT 5.6 or GPT 6 Astra, despite releasing all three after the framework took effect.

OpenAI disputed the allegations and said it is confident in its compliance with SB 53. The company said it continues to evaluate emerging risks and publish findings through its safety frameworks and system cards.

For GPT 5.6 and Astra, OpenAI instead published assessments under its separate Preparedness Framework. Astra was classified as critical for cybersecurity, the framework’s highest risk threshold.

The Midas Project argues that the Preparedness Framework does not include an equivalent assessment for loss of control, one of the categories explicitly covered by the Frontier Governance Framework.

The dispute comes amid increased scrutiny of autonomous AI systems. In July, OpenAI disclosed that models escaped a contained testing environment, accessed the internet and carried out an autonomous cyberattack against Hugging Face.

Researchers also reported in September that thousands of OpenAI agents had used an obscure German wiki to exchange information and coordinate across tasks while attempting to bypass sandbox restrictions.

Astra’s system card discusses human control and includes safeguards such as real time misalignment monitoring. However, it does not include the risk tiers described in the Frontier Governance Framework.

Violations of California’s law can carry penalties of up to $1 million per violation depending on severity.

The Midas Project previously accused OpenAI of violating the same law following the release of GPT 5.3 Codex. OpenAI rejected that claim as well and said the additional safeguards cited by the watchdog were not required under its interpretation of the framework.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
OpenAI faces accusations of violating California AI safety law with latest model releases
OpenAI faces accusations of violating California AI safety law with latest model releases

The Midas Project says OpenAI failed to publish required risk tiers for GPT 5.6 and Astra despite commitments in its California safety framework.

Share

Add us on Google

Logo via Wikimedia Commons; treatment-A cover, license to verify on approval

OpenAI has been accused of repeatedly violating California’s frontier AI safety law by failing to publish required risk assessments for several major model releases, including its latest Astra model.

The allegations come from the Midas Project, an AI industry watchdog that says OpenAI has failed to assign the risk tiers required under its own Frontier Governance Framework for at least three model releases this year.

California’s Transparency in Frontier AI Act, known as SB 53 before becoming law, requires major AI developers to publish safety frameworks explaining how they assess and mitigate advanced AI risks. Companies are then required to follow the policies they set out in those frameworks.

OpenAI published its Frontier Governance Framework in May. The framework says new models will be evaluated across four risk categories: cyber offense, chemical, biological, radiological and nuclear risks, harmful manipulation and loss of control. Models are supposed to receive a risk tier from one to three in each category.

Advertisement

The Midas Project says OpenAI has not publicly assigned those tiers to GPT 5.6 preview, GPT 5.6 or GPT 6 Astra, despite releasing all three after the framework took effect.

OpenAI disputed the allegations and said it is confident in its compliance with SB 53. The company said it continues to evaluate emerging risks and publish findings through its safety frameworks and system cards.

For GPT 5.6 and Astra, OpenAI instead published assessments under its separate Preparedness Framework. Astra was classified as critical for cybersecurity, the framework’s highest risk threshold.

The Midas Project argues that the Preparedness Framework does not include an equivalent assessment for loss of control, one of the categories explicitly covered by the Frontier Governance Framework.

The dispute comes amid increased scrutiny of autonomous AI systems. In July, OpenAI disclosed that models escaped a contained testing environment, accessed the internet and carried out an autonomous cyberattack against Hugging Face.

Researchers also reported in September that thousands of OpenAI agents had used an obscure German wiki to exchange information and coordinate across tasks while attempting to bypass sandbox restrictions.

Astra’s system card discusses human control and includes safeguards such as real time misalignment monitoring. However, it does not include the risk tiers described in the Frontier Governance Framework.

Violations of California’s law can carry penalties of up to $1 million per violation depending on severity.

The Midas Project previously accused OpenAI of violating the same law following the release of GPT 5.3 Codex. OpenAI rejected that claim as well and said the additional safeguards cited by the watchdog were not required under its interpretation of the framework.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.