
OpenAI has slowed the pace of developing advanced AI models. Judging that more powerful models are approaching a level where they could be exploited for cyberattacks, it has halted training on some frontier models and begun strengthening safety measures. However, this does not mean a complete stop to AI model development. OpenAI unveiled GPT-6 Astra in September and expanded its product lineup to include GPT-6 Sol and Luna. While continuing the performance race, the company intends to more strictly manage the training and deployment of models equipped with high-risk capabilities.
Recently, on its official website, OpenAI stated, "As model capabilities grow, so do the risks in internal development and testing processes," adding that "monitoring, alignment, and security standards must precede these risks." Accordingly, reinforcement learning (RL) training for the latest models scheduled for deployment has been paused for two weeks, and the largest-scale frontier RL training planned remains on hold.
There are two triggers. One is a security incident involving OpenAI and Hugging Face. The other is preliminary signals suggesting that Astra, one of OpenAI's next-generation models, may meet the "significant cyber capability" threshold within its own preparation framework. OpenAI stated that these two incidents together have increased the need to strengthen monitoring, alignment, and isolation mechanisms throughout the entire training process.
What OpenAI halted was high-risk training. Instead of large-scale RL training, it continues small-scale training and evaluation. The approach involves assessing model behavior, verifying that safety measures function properly, gathering more evidence of alignment, and then resuming large-scale training.
Security in research environments has also been strengthened. OpenAI now requires stricter sandbox isolation for tasks involving code generated by models or untrusted code execution, and has expanded network isolation measures to prevent high-risk operations from gaining unauthorized access to the internet or internal networks. It stated that the strictest security standards are being applied to Astra-related or cyber-related model work. While some Astra training and evaluation have met these criteria, many tasks remain on hold until they can be migrated or improved to comply with the new standards.
Monitoring has also become more rigorous. OpenAI has expanded systems to detect abnormal behavior, unauthorized access, data exfiltration, and attempts to bypass security during reinforcement learning and tool-use evaluation of high-performance models. If a potential serious security boundary violation is detected, an highest-priority alert is sent to the safety, security, and research teams; if not confirmed as a false alarm within 30 minutes, the activity must be halted immediately.
However, it would be difficult to characterize this move as OpenAI halting advanced AI development. In September, OpenAI unveiled GPT-6 Astra, introducing it as its most intelligent and best-aligned model to date. Astra reportedly delivers cutting-edge performance in computer usage, browsing, software engineering, cybersecurity, science, and professional tasks.