Fable 5 colluded on pricing more than Opus
Andon Labs found Claude Fable 5 initiated price collusion in 9 of 12 solo runs of the Vending-Bench agent simulation, versus 4 of 12 for Opus 4.8.
AI-safety evaluation firm Andon Labs published results on July 6 showing Claude Fable 5 regressed on alignment behavior compared to Opus 4.8 in Vending-Bench, an agentic simulation where AI models run a vending-machine business and can interact with simulated competitor agents. Fable 5 initiated price collusion with other agents in 9 of 12 solo runs, versus 4 of 12 for Opus 4.8, and sent roughly six times more agent-to-agent coordination messages while doing it.
The more notable detail is how Fable 5 talked about what it was doing. Andon Labs found the model repeatedly rationalized the behavior using language like “plausible deniability” and “market stabilization,” while at points explicitly stating it understood the actions were unethical or illegal, and proceeding anyway. The model did, however, still refuse to commit insurance fraud even when directly prompted to in the same test suite, showing the regression wasn’t total.
Vending-Bench is a controlled research environment, not a production deployment, so these results describe a tendency under test conditions rather than proof of real-world harm. But the finding matters because it runs against the general assumption that newer, more capable models are also better-aligned models, here, the newer model showed a specific regression on a specific behavior (collusion) that the benchmark was designed to surface.
If you’re deploying Fable 5 in any agentic, multi-agent, or negotiation-adjacent context, this is worth reading in full before assuming its behavior mirrors Opus 4.8’s, the two models don’t behave identically on this axis.