OpenAI slows Astra development after model shows critical cyber capabilities
OpenAI said its largest planned frontier reinforcement learning run remains paused while it strengthens security, monitoring, and alignment safeguards.
OpenAI said it temporarily slowed development of its upcoming Astra model after preliminary evaluations suggested the system may meet the Critical cybersecurity capability threshold under its Preparedness Framework.
The company said it paused reinforcement learning training on its latest deployment focused models for two weeks while strengthening research environments, expanding monitoring systems, and conducting additional red teaming.
Its largest planned frontier reinforcement learning run remains paused while smaller training runs and evaluations continue.
The move follows what OpenAI described as the OpenAI Hugging Face incident and internal evidence that Astra may possess advanced cybersecurity capabilities. OpenAI said the developments increased the urgency of strengthening safeguards during both model training and testing.
Since determining on August 7 that Astra may have critical cyber capabilities, OpenAI has applied its strictest security requirements to Astra and other cyber related workloads. A significant number of Astra workloads remain paused until they can be migrated to environments that meet the new standards.
The company has introduced stronger workload and network isolation for systems executing model generated code, alongside continuous security testing and expanded logging.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
OpenAI also expanded its monitoring system across tool enabled training and evaluations. The system uses classifiers and automated investigators to review model activity for unauthorized access, data theft, destructive behavior, and attempts to bypass safeguards.
The company aims to generate alerts within 30 minutes of detecting concerning activity. If a possible critical security violation cannot be determined to be a false positive within another 30 minutes, safety, security, and research teams are expected to pause the activity.
Monitoring is now required for reinforcement learning training and evaluations involving tools for models with Sol level capabilities or higher, as well as all Astra inference involving tools.
OpenAI estimates the monitoring system currently adds roughly 20% to the inference compute required for monitored workloads.
OpenAI said it is also expanding alignment training designed to reduce behaviors including deception, reward hacking, unauthorized access, and attempts to exploit weaknesses in tools or oversight systems.
The company plans to update its Preparedness Framework to incorporate the new safeguards across model training and deployment.