6,000 tries fail to hijack an AI assistant
Archive item — written before sources were shown.
A public challenge had 2,000+ people send 6,000+ emails trying to make an AI email assistant leak a secrets file. On Claude Opus 4.6, none succeeded.
Developer Fernando Irarrazaval ran HackMyClaw, a challenge to make an email assistant leak a planted secrets file through prompt injection. He reported more than 6,000 emails from over 2,000 people, including authority impersonation and multi-language social engineering, and zero successful extractions. The assistant ran on Claude Opus 4.6, which Anthropic trained to resist injection.
- 01What happened after 2,000 people tried to hack my AI assistantfernandoi.cl · primary
- 02This AI agent survived 6,000 hack attemptsdecrypt.co · reporting
