Logic of Logic
thursday, august 6, 2026 · the day's ai, attributed published by trilot llc · wyoming
brief safety

6,000 tries fail to hijack an AI assistant

A public challenge had 2,000+ people send 6,000+ emails trying to make an AI email assistant leak a secrets file. On Claude Opus 4.6, none succeeded.

Developer Fernando Irarrazaval ran HackMyClaw, a challenge to make an email assistant leak a planted secrets file through prompt injection. He reported more than 6,000 emails from over 2,000 people, including authority impersonation and multi-language social engineering, and zero successful extractions. The assistant ran on Claude Opus 4.6, which Anthropic trained to resist injection.

sources 2 cited
1 fernandoi.cl What happened after 2,000 people tried to hack my AI assistant 2 decrypt.co This AI agent survived 6,000 hack attempts
next