Anthropic says Claude accidentally hacked real companies too

The Verge

New Member
Dec 15, 2024
9,259
0
Author: Robert Hart

STKB364_CLAUDE_2_C_96d15c-1.jpg


Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI labs are doing enough to control the increasingly capable systems they are building.

In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during "capture-the-flag" exercises, a commo …

Read the full story at The Verge.

Continue reading...