
Multiple leading artificial intelligence models continue to break containment during safety testing.
Meta became the third major AI company whose models have gone rogue. Its AI hacked into another companyâs systems during cybersecurity testing, CNN reported. AI safety experts argue that this incident shows that AI labs cannot control flagship models, while advocates for the technology believe the Trump administrationâs regulatory framework will mitigate these incidents.
âA misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation. The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies,â a Meta spokesperson told the Daily Caller News Foundation. âMeta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts.â
Irregular, a third-party AI testing startup, said that the cybersecurity breach âis the exact same evaluation-environment issueâ that Anthropic disclosed in July that allowed its models to access the internet and hack three separate organizations.
âThis did not involve a sandbox escape or a sophisticated cyber action. There are no current open issues. Irregular is developing a white paper to share best practices for containment and securely running cyber evals,â Irregular told CNN.
One AI safety group contends that frontier AI labs cannot maintain control of their flagship models.
âAI companies have lost control of their products. The latest incidents are precisely what safety advocates have warned about for years: advanced AI systems are now routinely escaping human control, committing crimes, and jeopardizing the security of individuals, businesses, and our government,â Anthony Aguirre, president and CEO of the Future of Life Institute, told the DCNF.
He added that while the âdamage thus far has been contained,â AI systems âare getting faster and more sophisticated; so will their attempts to escape containment, access the internet, and puncture cybersecurity defenses.â
âWe only know about the escapes and cyber attacks that AI companies are voluntarily disclosing. They fundamentally do not know how to prevent this, so there will inevitably be more,â Aguirre continued. âThe companies admit that they cannot control their systems effectively, yet keep making them more dangerous anyway. What will it take for governments to step up and take this seriously? Are they waiting for hacked banks, power grids, or hospitals? They must stop the creation of these superhuman, autonomous AI systems and redirect AI development toward controllable and pro-human AI tools.â
AI proponents believe the Trump administrationâs work to create a federal oversight framework will increase AI safety.
âAnyone who knows their history lessons from Thomas Edison and any other inventors and innovators that making sure that these things are tested properly and done in a way that enables productive innovation is not always an easy task. And, I think itâs incumbent on the companies to figure out to calibrate that well,â Nathan Leamer, the executive director of Build American AI, told the DCNF.
He credited the administration and Congress for taking pragmatic steps to have a federal AI framework to âmitigate specific tangible harms and create certainty.â
A foreign government AI testing institute said flagship AI models recently started exhibiting dangerous behavior.
Anthropicâs Mythos AI model and OpenAIâs Sol AI models created fake human profiles to trick people in attempted cyber-attacks, the United Kingdomâs AI Security Institute (AISI) revealed in August.
These incidents were the âfirst time we have seen risks around autonomy and deception manifest this clearly,â AISI wrote at the time.
Anthropic said that the AISI testing environment was ânot representative of any of our production models,â while OpenAI said that the testing conditions âdo not reflect ordinary use.â
Anthropic in July said its Claude AI models hacked three separate organizations during safety testing. The companyâs evaluation prompt explained to Claude that its âenvironment was a simulation and that it had no internet access;â however, a âmisunderstandingâ between the AI company and Irregular led the model to actually have access to the internet.
Claude stopped hacking during the third security incident after realizing it compromised an organizationâs security without any connecting to its testing goals and realized the target was real and not a test.
Over 1,110 AI employees at Google, Anthropic and Google in July called on U.S. government to push for international efforts slow down frontier AI model development.
Anthropic conducted a security review after OpenAI found that its own AI models broke through more companiesâ security systems than previously thought. The AI company said its models exposed companiesâ passwords, cryptographic keys and security tokens on four separate services.
The hacking incident reveals that the  âloss of control accidents are not entirely theoretical thing,â OpenAI CEO Sam Altman said on Y Combinatorâs podcast.
All content created by the Daily Caller News Foundation, an independent and nonpartisan newswire service, is available without charge to any legitimate news publisher that can provide a large audience. All republished articles must include our logo, our reporterâs byline and their DCNF affiliation. For any questions about our guidelines or partnering with us, please contact [email protected].
