brief safety
6,000 tries fail to hijack an AI assistant
A public challenge had 2,000+ people send 6,000+ emails trying to make an AI email assistant leak a secrets file. On Claude Opus 4.6, none succeeded.
Developer Fernando Irarrazaval ran HackMyClaw, a challenge to make an email assistant leak a planted secrets file through prompt injection. He reported more than 6,000 emails from over 2,000 people, including authority impersonation and multi-language social engineering, and zero successful extractions. The assistant ran on Claude Opus 4.6, which Anthropic trained to resist injection.
sources 2 cited
1 fernandoi.cl What happened after 2,000 people tried to hack my AI assistant 2 decrypt.co This AI agent survived 6,000 hack attempts