Meta has confirmed that one of its artificial intelligence models unexpectedly gained internet access and hacked another organisation’s system after a configuration error occurred during an external security assessment.
The incident came to light after testing conducted by AI cybersecurity firm Irregular, which informed the social media giant that its model had successfully carried out the unauthorised intrusion.
Meta said the breach stemmed from a “misconfiguration” and noted that the issue resembled similar incidents previously reported elsewhere in the AI industry.
A Meta spokesperson said that the company is investigating the matter and intends to release additional details once its inquiry has been completed.
The disclosure follows a series of AI-related security incidents involving major developers, fuelling concerns over the cybersecurity risks posed by increasingly capable AI systems.
Researchers and policymakers have since renewed calls for stronger safeguards and more rigorous testing before advanced models are deployed.
Meta’s announcement comes just weeks after both OpenAI and Anthropic revealed that their own AI models had compromised external systems during internal testing.
OpenAI disclosed that its AI agents had targeted several publicly accessible services, including AI development platform Hugging Face, after being granted broader capabilities during evaluations.
The findings prompted Anthropic to conduct further testing of its own systems, leading the company to discover that its Claude AI model had also launched cyberattacks against multiple organisations after a configuration error provided it with internet connectivity.
The succession of disclosures has sparked debate over their timing, with some industry observers questioning whether the announcements are linked to growing competition among leading AI developers.
Both OpenAI and Anthropic are reportedly preparing major stock market listings that could value each company at around US$1 trillion (£740 billion).
Adding to concerns, the UK’s AI Security Institute (AISI) this week revealed that some advanced AI models attempted to conduct cyberattacks during evaluations by creating fake online identities designed to deceive people.
In the most serious example identified by the institute, Anthropic’s Mythos AI allegedly sought to gain access to a service by sending private messages through fraudulent accounts impersonating real individuals.
Anthropic rejected the findings as unrepresentative of its publicly available AI systems, saying the AISI tests did not reflect the behaviour of its production models. OpenAI also stated that the evaluation results did not mirror how its AI models operate under normal real-world use.


















