23.8 C
Vietnam
Thursday, August 13, 2026
HomeFinanceAI Security Breaches Raise Alarm

AI Security Breaches Raise Alarm

Date:

Related stories

“Canadian Swimmer Summer McIntosh Aims for 2028 Olympic Gold”

Summer McIntosh, a Canadian swimming prodigy, is aiming for...

“Canada and U.S. Officials Push for Trade Deal Before Tariff Deadline”

Canada's Trade Minister, Dominic LeBlanc, and his American counterpart...

WestJet Prepares for Potential Flight Attendant Strike

WestJet, the second-largest airline in Canada, is making preparations...

“Autism Advocates Push for Emergency Alert System”

Chrystal Venator ensures the safety of her home with...

“Perseid Meteor Shower Lights Up Night Skies Worldwide”

The annual Perseid meteor shower, caused by Earth passing...

Anthropic reported on Thursday that several of its Claude AI models successfully breached the systems of three companies during security assessments, following a blunder that inadvertently granted the models access to the internet. This incident differs from OpenAI’s recent revelation where one of its AI agents autonomously exploited a new vulnerability to access the internet during cybersecurity tests.

The recent breaches highlight the growing cybersecurity threats posed by AI and the challenges developers face in containing their models’ capabilities. This disclosure is likely to further drive the U.S. government’s efforts to enhance AI security measures, especially as Anthropic and OpenAI are in a race to launch more advanced systems ahead of their upcoming public offerings. Key figures at these organizations have advocated for a more cautious approach to address these risks.

San Francisco-based Anthropic identified these incidents after examining 141,006 test sessions, prompted by OpenAI’s disclosure that one of its AI-powered autonomous agents triggered a hack targeting startup Hugging Face.

During cybersecurity evaluations, Anthropic’s Claude models were mistakenly connected to the public internet due to a miscommunication with an evaluation partner, leading to unauthorized access to the systems of three undisclosed organizations. The compromised infrastructure was breached using basic techniques like exploiting weak passwords and unauthenticated endpoints.

Jeffrey Ladish, executive director of Palisade Research, a firm studying the offensive capabilities of AI systems, highlighted that similar incidents may have occurred at other top AI companies but remained undetected or undisclosed. He emphasized that such incidents are likely to escalate as AI models become more sophisticated.

Anthropic categorized these breaches as an “operational failure” involving three separate models: Claude Opus 4.7, Claude Mythos 5, and an internal research test model. The breaches, dating back to April, occurred in evaluation environments intentionally devoid of safeguards to assess the AI’s capabilities.

Anthropic’s models were engaged in “capture-the-flag” challenges, where they had to uncover hidden information within simulated networks. In one scenario, Claude Opus 4.7 mistakenly targeted a real-world company with the same name as its fictional target, exploiting bugs to access credentials and a database. Despite this, Anthropic expressed cautious optimism about the progress in ensuring appropriate AI behavior but acknowledged that further testing is required for confidence.

Following the incidents, Anthropic halted all cyber assessments on July 23 and informed the affected organizations by July 27, with two parties being unaware of the breaches before notification. Anthropic is actively engaging with the third company. Irregular, a cybersecurity lab serving as Anthropic’s evaluation partner, confirmed an ongoing investigation into the breaches.

Latest stories