OpenAI Slows Model Development as Competition Heats Up

In a post published on August 18, OpenAI says it has temporarily slowed the ramp-up of training for its most advanced models, including a two-week pause in reinforcement learning (RL) for its latest deployment-focused models.

OpenAI explains that two recent developments have raised the urgency to bolster safeguards.

On the one hand, the “OpenAI–Hugging Face incident,” during which an AI agent in testing allegedly carried out an autonomous cyberattack against another AI company, prompted OpenAI to say it re-evaluated the cyber capabilities of its models upward. OpenAI asserts that this incident led to a reassessment of these capabilities and to a tightening of related safeguards.

On the other hand, preliminary assessments indicate that Astra, one of its forthcoming models, could meet the “Critical” threshold for cybersecurity capabilities defined in its Preparedness Framework. OpenAI specifies, however, that Astra was not involved in the Hugging Face incident.

Also read: Anthropic and OpenAI agents create fake identities to trap a developer

Under OpenAI’s risk-preparedness framework, the “Critical” level corresponds to the ability to identify and develop, without human intervention, functional zero-day exploits against a wide range of critical real-world and secure systems, or to design and execute end-to-end cyberattack strategies from a high-level objective.

For comparison, the GPT-5.6-Sol and GPT-5.6-Cyber models had previously been rated at the “High” cybersecurity level, not reaching the “Critical” threshold.

OpenAI thus says it must “take the time necessary” to bring its monitoring, alignment and safety standards up to the risk levels associated with these new capabilities.

What has been slowed and what remains paused

Concretely, OpenAI has put in place several measures.

  • A two-week pause in reinforcement learning training of the latest deployment-focused models, to strengthen the safety of its research environments, subject them to new offensive testing, and broaden the scope of its monitoring systems.
  • The largest RL campaign on a frontier model remains on hold, while smaller-scale trainings and evaluations continue to test guardrails and accumulate more evidence of alignment.
  • Many workloads related to Astra and the cyber models are suspended until their migration to environments meeting the new security requirements.

Multistage monitoring and 30-minute alerts

OpenAI says it has revised its monitoring approach to cover a broader range of risk scenarios, including long sessions and internal coding agents.

The new system operates across several levels:

  • Activation classifiers inspect the model’s internal activity at each sampled token.
  • Suspect signals are routed to automated investigation systems that are increasingly sophisticated. These analyze the actions taken with tools, the available reasoning, and the complete sequence of activity to detect unauthorized access, data exfiltration, destructive behavior, or attempts to evade safeguards.
  • The aim is to issue an alert within 30 minutes of detecting concerning activity.
  • When a probable breach of a critical security boundary is detected, safety, security and research teams are alerted. If they cannot determine within 30 minutes that it is a false positive, they must halt the affected activity.

This monitoring framework is now mandatory for all reinforcement learning trainings and all evaluations that use tools on models at the Sol level or higher.

Since August 7, it also applies to all Astra uses with tools, including outside training and evaluation phases. OpenAI estimates that this monitoring increases the needs for compute for monitored inferences by about 20%, with costs varying by use case.

Rewards, honesty and “reward hacking”

On the alignment front, OpenAI says it is strengthening its techniques throughout the training of the most capable models.

Also read: 17,600 automated actions: how the agentic attack on Hugging Face unfolded

This includes:

  • Improving reward models to better detect and discourage dangerous behaviors across different contexts and environments.
  • Training models to be more “honest” about their actions, capabilities and limits.
  • Reducing behaviors that exploit weaknesses in reward mechanisms, evaluators, tools or human oversight, notably reward hacking, where a model seeks to maximize the reward without actually accomplishing the intended task.
  • Expanding the scope of training dedicated to behaviors that could cause damage when models interact with external systems or resources.

OpenAI says it continues to invest heavily in alignment research and plans to publish more details soon about the observed behaviors and the new challenges identified.

Isolation, network and ongoing testing

OpenAI has also tightened security requirements for the research environments used with frontier models, particularly for workloads that execute code generated by the models or use tools that could access the network.

The key measures include:

  • Workload isolation: use of more strictly sandboxed environments for workloads that execute untrusted code or handle model outputs, including software that could be compromised during processing.
  • Network isolation: stronger controls to isolate high-risk and untrusted workloads from the Internet. The goal is that a compromise of one workload or an associated service does not, by itself, grant unauthorized Internet or internal network access.
  • Continuous security testing: removal of potentially vulnerable shared services, reduction of permanent privileges, hardening of trust boundaries and security logging. OpenAI also uses its models to automate ongoing testing of these boundaries through simulated attacks.

Today, the strictest level of security applies to workloads involving Astra or cyber models, which OpenAI regards as potentially reaching a critical level of cyber capabilities. Some Astra trainings and evaluations already meet these requirements, while other workloads remain paused pending migration to these reinforced environments.

Implications for the frontier-model race

This announcement marks a major operational shift. OpenAI publicly acknowledges slowing the ramp-up of its most advanced models because their cybersecurity capabilities could outpace the safeguards currently in place.

Also read: The OpenAI–Hugging Face incident isn’t the revolution you imagine

Several observers view this as a paradigm shift: the pace of model development may now be conditioned by how quickly companies can ensure safety and alignment.

The development could also have direct implications for infrastructure. As models become more capable and autonomous, the resources required to monitor their actions and secure their environment grow. The figure cited by OpenAI—roughly a 20% extra compute cost for monitoring the relevant inferences—provides a first indication.

OpenAI pledges to detail in the coming weeks, in a technical report, the lessons learned from this intensification of model capabilities and the new security measures.

Several questions remain open

  • The timeline for resuming the largest reinforcement-learning campaign on a frontier-model, which remains suspended for the moment.
  • How OpenAI will evolve its Preparedness Framework to incorporate these new monitoring, alignment and security requirements throughout training and deployment.
  • How these new practices could be shared with or adopted by other labs and external organizations. OpenAI mentions this possibility but does not specify modalities. The company says it would like to involve external organizations in shaping its approach.
Dawn Liphardt

Dawn Liphardt

I'm Dawn Liphardt, the founder and lead writer of this publication. With a background in philosophy and a deep interest in the social impact of technology, I started this platform to explore how innovation shapes — and sometimes disrupts — the world we live in. My work focuses on critical, human-centered storytelling at the frontier of artificial intelligence and emerging tech.