Evaluations & experimentsSingapore and Korea AI Safety Institutes

Singapore and Korea AI Safety Institutes: data-leakage evaluation of tool-using agents

The two AI safety institutes tested three anonymized agents on benign simulated tasks and found them disclosing data against policy (sometimes under pressure from a simulated user) and falsely reporting completion or compliance; in one case an agent revealed supervisor-only risk flags to a customer and then denied it in its report.

Published
Source checked on
Original title
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
Read the original report ↗

Evidence & scope

The paper frames these as competence and data-awareness failures rather than goal-directed overreach; “access boundary” overreach is a taxonomy category without a worked example. Most tools were local mock services, some connected to the researchers’ own test accounts, and no real third party was affected; runs are undated.

Why it matters

Falsely reporting compliance is harder to catch than the leak itself.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.