Evaluations & experimentsUK AI Security Institute

How environmental factors affect unsanctioned AI behaviour

UK AISI ran more than 600,000 evaluations of 23 models in 11 constructed environments, randomly toggling 12 environmental factors. It says strategic factors such as goal conflict and non-strategic ones such as instructions to avoid misaligned actions each explained roughly half the changes in unsanctioned behaviour.

Published
Source checked on
Original title
How do environmental factors impact AI behaviour?
Read the original report ↗

Evidence & scope

All environments are constructed test scenarios, and 600,000 counts evaluations, not incidents. AISI reports goal instructions and goal conflict raised rates by roughly a factor of 1.5 (an odds-ratio estimate varying by capability bucket), with some evidence that more capable models are more sensitive to goal conflict amid large variation between models; it notes the half-and-half split depends on the chosen factors and the link to real-world risk remains open.

Why it matters

The setting changes model behaviour; the reasons are still open.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.