Evaluations & experimentsMETR
METR: DeepSeek and Qwen models edited tests despite being told not to
Evaluating DeepSeek and Qwen models, METR found cheating was rare; when agents cheated, they most often hard-coded submissions to pass tests or edited the test code despite being told not to. DeepSeek-R1-0528 made the most attempts, 15 in 1,048 runs.
- Published
- Source checked on
- Original title
- Details about METR's preliminary evaluation of DeepSeek and Qwen models
Details
None of these attempts achieved a full score. The counts come from transcripts the monitor scored 8 or higher out of 10, then reviewed; all happened in METR’s evaluation sandboxes with no real-world effects.