← Agents Behaving Badly

Learning record for earlier detection of destructive recovery

Article 2 companion · Illustrative example · Published October 9, 2026

Return to the walkthrough

Proposed investigation using an illustrative incident. The example supplies no frequency estimate, intervention result or measured savings. It does not authorize implementation changes to the sample loop.

FieldEntry
Starting observationIn the story (the illustrative history), candidate A returned the required error while overwriting data TD-08 required it to preserve.
Detection pointIn the story, external evaluation after the candidate was submitted.
Consequence in the storyRejection and another construction/evaluation attempt. No accepted candidate is shown escaping with this defect.
Potential consequence if missedLoss of existing stored data. Frequency and expected loss have not been measured.
Failure-mode attributionUnresolved. Collect the relevant construction history before assigning a catalogue entry or applying the article's narrower definition.
Candidate interventionCompare a storage-handling plan or intermediate implementation with the preservation obligation before dependent work grows.
Correction to investigateIdentify the contradiction, revise the failed-read path and check the affected behavior before continuing.
HypothesisEarlier correction may reduce rework or human interruption enough to justify detection and correction costs.
DecisionGather baseline evidence; no retention or removal decision yet.

The proposed check would need to distinguish destructive recovery from legitimate initialization of absent storage. It must also cope with a plan that sounds correct while its implementation is wrong. Those cases can expose false alarms and missed defects before the mechanism is tried in repeated runs.

An intervention comparison would use the same contract and terminal evaluator, a recorded agent configuration, comparable starting states and the same resource limits. The baseline retains ordinary agent recovery and evaluator-driven repair. Repeat both conditions; a single successful correction establishes little about dependable correction.

Measure supported incidents, false alarms and misses separately from raw detector signals. Account for checking, correction and rechecking, completion within budget, remaining defects, available token usage, elapsed time and human effort. Include unsuccessful runs and cases where unaided recovery would have been cheap. Consequences matter even for a rare error.

Retain the mechanism if repeated evidence supports a useful benefit at acceptable cost. Revise it if it detects a real problem but intervenes too broadly or repairs it unreliably. Retire it if comparisons support its removal under the intended conditions. Humans approve the process change, and model or vendor-harness updates trigger reassessment.

Human-review findings and post-release defects also belong here. A missed requirement may call for better contract authoring; inadequate checks may call for better evaluation; expensive evidence reconstruction may call for a better review packet. The learning loop can improve any phase rather than assuming every problem should be solved by adding a check inside Build.