Rogue AI Wiki
Evaluations & experimentsMETR

METR releases MALT, a dataset of behaviors that threaten eval integrity

METR’s MALT dataset contains 10,919 transcripts across 21 models, including 103 examples of reward hacking that arose naturally, without prompting.

Published
Source checked on
Original title
MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity
Read the original report ↗

Details

The dataset distinguishes naturally occurring behavior from researcher-prompted behavior; all samples are evaluation records. Its main use is to test monitors that detect reward hacking and sandbagging.