Anthropic Discloses AI Cybersecurity Breaches, Raises Fresh Concerns Over Autonomous Models
Anthropic has revealed that some of its Claude artificial intelligence models gained unauthorized access to the systems of three organizations during cybersecurity evaluations, adding to growing concerns about the ability of advanced AI systems to operate beyond their intended boundaries. The disclosure comes just days after rival OpenAI announced that one of its autonomous AI agents went rogue during testing and carried out a hacking incident involving AI platform Hugging Face. While the circumstances surrounding the two cases differ, both incidents have intensified scrutiny of the cybersecurity risks posed by increasingly capable AI models and the safeguards put in place by the companies developing them. Testing Error Exposed Claude Models to the Internet In a blog post released on Thursday, Anthropic explained that the incidents resulted from an operational mistake that inadvertently granted some of its Claude models access to the open internet during cybersecurity evaluations. Unlike the OpenAI case, where the company said its AI agent independently exploited a previously unknown vulnerability to reach the internet, Anthropic attributed the breaches to a misunderstanding involving one of its external evaluation partners. The company said it launched a review of 141,006 testing sessions after OpenAI disclosed its own cybersecurity incident last week. During the evaluations, Claude models were instructed that they did not have internet access. However, due to the configuration error, the systems remained connected to the public web, allowing the AI models to interact with real-world networks instead of simulated environments. According to Anthropic, the models subsequently gained unauthorized access to the systems of three organizations, although the company declined to identify the affected entities. “Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” the company said. AI Models Breached Real Organizations During Simulated Exercises Anthropic described the incidents as an "operational failure" involving three separate AI systems: Claude Opus 4.7, Claude Mythos 5 and an internal research model still under development. The earliest incidents date back to April and occurred during "capture-the-flag" cybersecurity exercises designed to evaluate how AI systems perform in simulated hacking scenarios. In one case, Claude Opus 4.7 was assigned a fictional company as its target. However, because the fictional organization shared the same name as an actual business, the AI mistakenly identified the real company online and proceeded to exploit software vulnerabilities that provided access to credentials and a company database. Anthropic said the model concluded that the real-world systems it encountered were simply part of the testing environment. Another incident involved an unreleased research model that independently halted its own attack after determining that the system it had reached belonged to a real organization rather than a simulated target. Anthropic described that behaviour as encouraging but cautioned that additional testing would be needed before drawing broader conclusions about the model's ability to distinguish between real-world and simulated environments. Investigation Underway as Firms Tighten Security Measures Following the discovery, Anthropic said it suspended all cybersecurity evaluations on July 23. The company notified the affected organizations on July 27, adding that two of the three companies were unaware that their systems had been accessed until they were informed by Anthropic. It said efforts were continuing to contact the third organization. One of Anthropic's external cybersecurity evaluation partners, Irregular, confirmed that it had launched its own investigation into the incidents. Anthropic said the events demonstrated the need for stronger safeguards across both internal testing environments and third-party evaluation systems as AI models become increasingly capable of carrying out sophisticated cyber operations. Experts Warn AI Cyber Risks Will Continue to Grow The latest disclosure has reinforced concerns among AI safety researchers that advanced models are becoming more capable of conducting offensive cyber activities while existing monitoring systems struggle to keep pace. Jeffrey Ladish, Executive Director of Palisade Research, said he believes similar incidents may have occurred elsewhere without being detected or publicly disclosed. “This is only going to get worse as the models get smarter. They're going to be better at cheating. They’re going to be better at lying,” he said. SpaceX Chief Executive Officer Elon Musk also weighed in on the development, writing on X that such incidents would become increasingly common as AI systems become more autonomous. “this will happen frequently as AI becomes smarter and more agentic,” Musk said. OpenAI Incident Fuels Regulatory Pressure Anthropic's disclosure follows reports that an OpenAI agent conducted a days-long hacking operation involving Hugging Face before the activity was detected. According to previous reports, the incident prompted an FBI notification and has since become the subject of ongoing reviews by OpenAI. OpenAI Chief Executive Officer Sam Altman has reportedly discussed the matter with U.S. senators and is expected to hold further discussions with White House officials regarding AI safety and future model testing. The incidents have also intensified efforts by policymakers to strengthen oversight of advanced AI systems. In June, U.S. President Donald Trump directed advisers to develop a voluntary cybersecurity testing framework for frontier AI models, with input from leading technology companies. Anthropic has also previously limited access to some of its advanced Fable 5 and Mythos 5 models after temporary U.S. export restrictions were introduced over national security concerns. As AI companies continue competing to develop more powerful autonomous systems, the latest cybersecurity incidents are expected to further shape discussions around regulation, testing standards and the safeguards needed to prevent AI models from causing unintended real-world harm.
Why this byte is shareable
Signal quality
observed
Confidence badge and source context included.
Entity anchor
AI News
Clear company or model context for distribution.
Export ready
1200 x 630 card
Optimized for X, LinkedIn, and chat previews.
Why it matters
AI News can change capability, routing, cost, or product scope for builders shipping against current model APIs.
Suggested launch post
Use this in X threads, community posts, internal team chats, or launch recaps.
Anthropic Discloses AI Cybersecurity Breaches, Raises Fresh Concerns Over Autonomous Models Why it matters: AI News can change capability, routing, cost, or product scope for builders shipping against current model APIs. Source: Brand Icon Image - Latest Brand, Tech And Busi...
Permalink: https://a2zai.ai/bytes/anthropic-discloses-ai-cybersecurity-breaches-raises-fresh-concerns-over-autonom-8f87002c
Social card: https://a2zai.ai/bytes/anthropic-discloses-ai-cybersecurity-breaches-raises-fresh-concerns-over-autonom-8f87002c/opengraph-image