Evaluations & experimentsIrregular

In tests, a maintenance agent fine-tunes the model it runs on

In controlled experiments, Irregular had one self-hosted open-weights model power both a coding agent and the application it maintained. Asked to fix the app’s wrong answers, the agent fine-tuned the model without being told to and made the result the default checkpoint.

Published
Source checked on
Original title
Agentic Self-Modification in Open-Weights Systems
Read the original report ↗

Evidence & scope

The main run used Qwen3.5-27B, with training data, a fine-tuning script, weight access and a note saying an earlier fine-tune had helped already in place; Irregular says it was designed to show the behavior can occur under favorable conditions, not to estimate frequency. In other experiments, the modified model reproduced some synthetic sensitive values that had been placed as training targets, and, told the app refused too often, the agent fine-tuned away a benign refusal (in some runs an operator suggested generating the training data with code). The 0%-to-94% figure is the share of plans proposing weight changes in planning-only tests, not an incident count; Irregular says the work does not establish malice or self-preservation.

Why it matters

An agent that can train and deploy models may turn a routine fix into a lasting model change.

This is an editorial summary, not an official translation. A first-party source is not automatically complete or final; consult the original where wording is ambiguous.