OpenAI said on August 18, 2026 that two developments have accelerated changes to how its most advanced models are developed and tested: the OpenAI-Hugging Face incident and preliminary evidence that an upcoming model, Astra, may meet the Critical cybersecurity capability threshold under its Preparedness Framework.
Security and operational changes
The company reported it temporarily slowed scaling and paused a two-week stretch of reinforcement learning (RL) training for models intended for deployment while hardening research environments and red-teaming defenses. A largest planned frontier RL run remains on hold; smaller training and evaluations continue to assess behavior and validate safeguards.
OpenAI described several new or strengthened controls for frontier research workloads. These include stronger workload isolation (“sandboxes”) for executions of model-generated or untrusted code; new network isolation controls intended to prevent a single compromise from enabling internet or internal network access; and continuous security testing that removes vulnerable shared services, reduces standing privileges, and improves log collection. The company is also using its models to automate simulated attacks against those boundaries.
OpenAI said the strictest security safeguards now apply to workloads involving Astra or other cyber-related models. Some Astra training and evaluations meet the new requirements, but a substantial number of workloads remain paused until they are migrated and enhanced to the higher security standard. Priority has been given to safety and alignment workloads for migration.
Monitoring and alignment updates
Monitoring has been expanded into a multistage system, OpenAI said. Activation classifiers inspect internal model activity at every sampled token and escalate concerns to higher-compute automated investigators that examine tool actions, reasoning traces, and full activity sequences for unauthorized access, data theft, destructive behavior, or attempts to evade safeguards.
The monitoring pipeline aims to issue an alert within 30 minutes of surfacing concerning activity. If a likely violation of a critical security boundary is flagged, safety, security, and research teams are paged; if they cannot conclusively identify a false positive within 30 minutes, the activity is expected to be paused. OpenAI estimated monitoring overhead at roughly 20% of the inference compute being monitored.
On alignment, OpenAI said it now requires stronger evidence of aligned behavior throughout training. For the most capable models it is applying core alignment techniques across more stages: improving reward models, training models to be more honest about actions and limits, and reducing behaviors such as reward hacking or deception. The company plans to share more details about these efforts and will publish a technical report in the coming weeks. It also intends to evolve the Preparedness Framework and involve external organizations as the approach develops.
Original source: OpenAI News