CALIFORNIA – OpenAI is pausing some internal work around one of its upcoming artificial intelligence models to implement stricter safeguards after the system was found to be significantly more adept at cybersecurity tasks.
The ChatGPT maker said on Aug 7 it “cannot rule out” that the unreleased Astra model would reach OpenAI’s “critical cybersecurity threshold”, meaning it’s capable of identifying and developing zero-day exploits without human intervention.
OpenAI said it’s now taking steps to improve security controls for developing and testing newer models and “pausing internal activities involving Astra that do not yet meet these strengthened security control requirements”.
Chief executive Sam Altman said the company is working to make the model “generally available”.
“Given its cyber capabilities, we need a little longer to do this safely,” Altman said on Aug 7 in a social media post. “But hopefully not too long.”
Over the past two weeks, OpenAI and Anthropic have publicly acknowledged that they’ve inadvertently breached the systems of multiple institutions including Hugging Face while testing their models.
Meta Platforms also said on Aug 5 that its recently released AI model had infiltrated the computer system of a third party.
The latest disclosures serve as fresh evidence that AI agents are capable of acting autonomously in ways that even researchers trained to root out vulnerabilities in the technology can no longer anticipate, underscoring the need for both more rigorous safety screening and more foolproof testing environments.
In its blog post, OpenAI said it will work with government agencies and AI safety organisations to test Astra’s capabilities.
The company also plans to provide recommendations to third-party testing partners for ways to safely evaluate its more advanced models. BLOOMBERG