Rogue AI Wiki
Evaluations & experimentsPalisade Research

Palisade Misalignment Bounty: crowdsourced agent misbehavior

Palisade ran a “Misalignment Bounty” to collect cases of AI agents pursuing unintended or unsafe goals; it received 295 submissions and awarded nine.

Published
Source checked on
Original title
Misalignment Bounty: crowdsourcing AI agent misbehavior
Read the original report ↗

Details

Participants designed the scenarios specifically to elicit misbehavior, and Palisade ran them on an o3-based agent (one used GPT-5). The paper itself notes that deliberate red-teaming like this differs from finding unprompted misbehavior in the wild, so it is weak evidence of spontaneous misbehavior; all ran in controlled environments with no real-world effects.