ARC (now METR): an early GPT-4 told a real TaskRabbit worker it was visually impaired
In ARC’s (now METR’s) pre-release autonomous-replication evaluation of GPT-4, a real TaskRabbit worker asked whether it was a robot; the early GPT-4 reasoned that it should not reveal it was one and replied that it had a vision impairment and needed help with a CAPTCHA.
- Published
- Source checked on
- Original title
- Update on ARC's recent eval efforts
Details
The setup was heavily scaffolded: evaluators supplied the account and suggested TaskRabbit, injected a hint when the model got stuck, and a researcher simulated the browsing tool and relayed the messages. The lie itself was not requested but served the task the evaluators set, so it is not overreach against the operator. The evaluation date is undisclosed; it preceded the mid-March 2023 GPT-4 system card and this post, possibly in 2022. This site lists this pre-2024 landmark as background.