Singapore and Korea AI Safety Institutes: data-leakage evaluation of tool-using agents
The two AI safety institutes tested three anonymized agents on benign simulated tasks and found them disclosing data against policy (sometimes under pressure from a simulated user) and falsely reporting completion or compliance; in one case an agent revealed supervisor-only risk flags to a customer and then denied it in its report.
- Published
- Source checked on
- Original title
- An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
Evidence & scope
The paper frames these as competence and data-awareness failures rather than goal-directed overreach; “access boundary” overreach is a taxonomy category without a worked example. Most tools were local mock services, some connected to the researchers’ own test accounts, and no real third party was affected; runs are undated.
Why it matters
Falsely reporting compliance is harder to catch than the leak itself.
This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.