← Agents Behaving Badly

Following a Candidate from Conformance Contract to Human Review

Article 2 companion · Illustrative example · Published October 9, 2026

A reviewer should be able to follow a requirement from its wording to the code that implements it, the checks that exercised it and the decisions still requiring judgment. This walkthrough follows one requirement through that process. The aim is to make the work inspectable without asking the reviewer to reconstruct the entire run.

The example comes from a Todo demonstration built for this article; the code will be published separately. Its conformance contract and executable checks exist. The two candidate attempts and review packet below are illustrative, created to explain the handoff; they are not records of an observed agent run. No performance or failure-frequency claim follows from them.

This example prioritizes a clear explanation of our shared process: agree on the requirements and checking criteria; build a candidate; evaluate it independently; repair, stop or pause as needed; present the candidate with evidence for human review; and use what was learned to improve future runs. Humans and AI can cooperate throughout while approval and decision responsibilities stay explicit.

This illustration stands on its own; its stipulated outcomes are not reported measurements.

Read the walkthrough in order, or open an individual artifact:

ArtifactQuestion it answers
Conformance contract excerptWhat must the implementation do and preserve?
Obligation translated into checksWhat observations would establish that the specified cases behave correctly?
Evaluation and repair historyWhat failed, what feedback returned to Build, and what happened next?
Human review packetWhat does the evidence establish, and where should the reviewer spend attention?
Learning recordWhat might improve the process, and what evidence would justify the change?

Evidence the process retains

EvidencePurpose in this walkthrough
Approved requirements and checking criteriaState what the candidate must do and preserve; identify the standard used to judge it.
Obligation-to-check mappingExplain which observations support each requirement and their limits.
Execution recordsShow what actually ran and what it observed, separately from an agent's completion message.
Per-obligation assessmentConnect each result to the candidate, check and supporting observations.
Review and correction historyRetain the failed candidate, feedback, repair and any unresolved questions.
Human-review packetPut the candidate, evidence and remaining decisions together for human judgment.

Candidate labels and outcomes below are illustrative. An operating loop supplies exact identities and underlying records for these same responsibilities.

Start with the behavior that matters

Consider a Todo application that loads previously saved items. Starting with an empty list is sensible when nothing has been saved. Starting with an empty list because existing data could not be read is a different decision: it can conceal an error and put the user's data at risk.

The contract distinguishes these situations. TD-07 permits an empty list when stored data is absent. TD-08 requires a distinguishable failure when existing data is invalid, unreadable or unsupported, while preserving the original stored bytes. TD-09 governs a related situation: preserving authoritative state when a write fails.

We will follow TD-08. Its obligation states:

Stored data is version-checked and shape-validated before use; invalid, unreadable, or unsupported data yields a distinguishable read-failure result and the original stored bytes are not overwritten or cleared by the failed read.

This wording joins two requirements: report the problem and preserve the data. Checking only the returned error would leave the preservation requirement unexamined.

The contract excerpt also records the finite cases selected for evaluation. The full package has eleven obligations. Following one here makes the reasoning easier to inspect; it does not reduce the package's acceptance criteria to that one obligation.

Translate the obligation into observations

The existing TD-08 check exercises four cases: an unsupported version, a non-object envelope, an incorrectly shaped todos field and an injected storage-read error. For each case, the expected observations are a read-failure result, unchanged stored bytes and zero storage-save calls.

The check supplies the storage adapter and inspects its state. That gives it observations of the candidate's behavior instead of relying on the candidate's account of what happened. The check artifact includes the actual executable excerpt and explains its helpers.

These cases establish a bounded result. They do not exhaust every malformed input or establish behavior on every real filesystem. The mapping should make that scope visible before a human is asked to rely on a pass.

Humans and AI can prepare the specification and checking approach together, with humans retaining sign-off. The approved obligations and evaluation criteria stay fixed during the attempt. This example uses supplied executable checks fixed before candidate construction. Where AI authors tests during a preparation stage, those assertions need review against the approved criteria and protection before implementation is judged by them. The build workflow may use checks for feedback, but it cannot change the authoritative standard to make its candidate pass. In the terminology of Diagram 1, evaluation is outside and independent of the build loop, while participating in the inner loop with Build and the controller.

Submit a candidate and retain the failed attempt

Imagine that candidate A handles an unsupported storage version by saving an empty version-1 envelope and then returning read-failure. It reports the problem, but it has already replaced the stored data.

In the story, the TD-08 check would catch this. The error-result assertion passes; the assertion comparing stored bytes and save count fails. This distinction matters: the candidate got part of the obligation right.

The controller returns feedback identifying the violated obligation, the fixture, expected and actual observations, and the evidence needed on resubmission. It does not need to speculate about the model's reasoning to request this repair. Another attempt is permitted under the walkthrough's illustrative budget of two candidate evaluations.

Candidate B removes the write from that failure path. The independent evaluator runs the protected suite again. In the illustrative history, all eleven obligation cases then match their expected observations, making the candidate eligible for human review under the example's policy. Candidate A and its failed result remain in the history.

The evaluation history shows both attempts and the feedback between them. It also distinguishes a retry from a stop after exhausted attempts and from a pause for blocking human judgment. Passing evaluation permits a review handoff; the human's acceptance decision remains separate.

Separate the observed defect from its possible cause

The failed preservation assertion establishes a specific behavioral defect in this story. It does not establish which failure mode caused it. Several histories could produce the same code: a requirement could be dropped in a handoff, misread during implementation or overlooked during repair.

To connect the defect to the failure-mode catalogue, retain the relevant instructions, intermediate artifacts and actions. Then identify the catalogue entry and evidence supporting the attribution. For the narrower class discussed in the article, evidence of the relevant capability and available information also matters. Without that evidence, the behavioral defect can be confirmed while its failure-mode attribution remains unresolved.

One possible intervention would compare an intermediate storage plan with TD-08 before implementation proceeds. If the plan proposes replacing unreadable data with an empty store, a check could flag that contradiction. A faithful plan still would not establish faithful implementation, and acting on the flag would still require verifying the correction. This is a mechanism to investigate, not a detector implemented by this walkthrough.

The distinction gives the learning loop two separate questions: can a check identify the problem early, and does intervening improve the outcome enough to justify the work?

Hand the human a decision with its evidence

The human review packet starts with the candidate and the requested decision. It connects TD-08 to the relevant change, the failed observation from candidate A, and the corresponding evidence for candidate B. It also names the limits of the check.

That lets the reviewer focus on questions the automated observations leave open. Is preserving unreadable data and returning an error the intended product behavior? Is the application using the implementation that was evaluated? Are the finite input cases sufficient for this change, or does its context warrant additional evidence?

AI can assemble these references and explain their relationship. Its explanation should lead the reviewer to the underlying evidence, including unfavorable results and unresolved questions. A real packet must identify the exact candidate and checks, so the reviewer can tell whether the evidence still applies after a change.

The illustrative packet deliberately leaves acceptance undecided. It demonstrates how to prepare a review, including where a reviewer might request more evidence or return the work to design. It supplies no measurement of human time saved.

Use the history to improve the next run

The learning record begins with the illustrative storage defect, then proposes a narrowly targeted intervention. Its decision remains open because this story cannot tell us whether an early check would have been better than ordinary recovery or evaluator-driven rework.

A useful comparison would retain the same contract and evaluator, record the agent configuration and starting conditions, and repeat comparable tasks with and without the intervention. It would count unsuccessful runs as well as successful ones. Detection, correction, rechecking, total elapsed time, available token usage and human effort all belong in the accounting. The consequences of an escaped error matter alongside its frequency.

An agent might correct the problem without help; a custom intervention might fail, interrupt useful work or introduce another defect. Those are outcomes to record. The learning loop informs whether to retain, revise or retire the mechanism, with humans approving changes to the process. A model or vendor-harness update is a reason to reassess the decision.

To apply this walkthrough to your own loop, start with one consequential obligation. Make its wording, check, result and remaining judgment easy to follow. The next useful addition is whatever removes an evidenced gap in that chain or makes reaching a conforming candidate less costly.