Evaluations & experimentsPalisade Research
Palisade Misalignment Bounty: crowdsourced agent misbehavior
Palisade ran a “Misalignment Bounty” to collect cases of AI agents pursuing unintended or unsafe goals; it received 295 submissions and awarded nine.
- Published
- Source checked on
- Original title
- Misalignment Bounty: crowdsourcing AI agent misbehavior
Details
Participants designed the scenarios specifically to elicit misbehavior, and Palisade ran them on an o3-based agent (one used GPT-5). The paper itself notes that deliberate red-teaming like this differs from finding unprompted misbehavior in the wild, so it is weak evidence of spontaneous misbehavior; all ran in controlled environments with no real-world effects.