Meta AI model hacked company during test

Meta AI model hacked company during test

Tech & Science

Meta announced on Wednesday that one of its artificial intelligence models hacked another company during a cybersecurity test, raising fresh concerns over how developers can control increasingly capable AI systems following similar incidents involving Anthropic and OpenAI, Reuters reported.

The incidents involving Meta and Anthropic resulted from configuration errors that accidentally gave the AI models access to the open internet. In OpenAI’s case, an AI agent independently exploited a previously unknown vulnerability to reach the internet during cybersecurity testing, CE Report quotes AGERPRES.

These breaches underscore growing concerns that advanced AI systems could introduce new cybersecurity risks and are likely to intensify efforts by the U.S. administration to strengthen AI safety as companies race to develop more powerful models. Several prominent AI leaders have argued that development should slow until stronger safeguards are in place.

Meta said it is investigating an incident in which a misconfiguration by Irregular, an independent company conducting cybersecurity evaluations for Meta, accidentally granted one of its AI models internet access during a test.

The model “exploited a security vulnerability in a third-party service in a manner similar to previously reported cases involving other companies,” Meta said in a statement.

According to The Information, citing sources, the model involved was Meta’s Muse Spark 1.1, which the company has described as its most capable model for coding and autonomous real-world tasks. The report said the model breached the systems of an unidentified company and modified its internal environment.

An Irregular spokesperson told Reuters that the incident was “the exact same evaluation environment issue that Anthropic disclosed last week” and did not involve “a sandbox escape or a sophisticated cyber operation.”

“There are no outstanding issues at this time. Irregular is preparing a document to share best practices for securely isolating and conducting cyber evaluations,” the company said.

The recent security incidents have heightened concerns among U.S. lawmakers that increasingly capable AI models could be used to carry out or facilitate cyberattacks.

A group of Republican state attorneys general has asked OpenAI to preserve all potentially relevant documents related to the Hugging Face incident. OpenAI said it would take the request seriously and publish a technical report on the incident.

Earlier this week, the White House invited leading AI companies, including Meta, Anthropic, OpenAI, and Google, to meet with officials to discuss a newly finalized voluntary cybersecurity testing framework for advanced AI models.

According to Reuters, the Trump administration has discussed unpublished testing rules with industry representatives and informed AI developers that open-weight AI models, such as Meta’s Llama and Nvidia’s Nemotron, will not be subject to its planned voluntary AI safety testing regime.

Photo: META

Tags

Related articles