Illustrative evaluation and repair history
Article 2 companion · Illustrative example · Published October 9, 2026
Illustration, not a recorded run. Candidates A and B are teaching labels. No native agent was run to produce this history, and there are no candidate commits, measured costs or raw run logs behind these labels. The TD-08 obligation and executable assertions come from the existing sample package.
For the story, assume the controller permits at most two candidate evaluations and holds the contract and authoritative checks fixed throughout. Assume the other ten obligation cases match their expected observations in both attempts. These assumptions let us follow one failure without pretending that TD-08 alone qualifies the candidate.
| Attempt | Illustrative candidate behavior | Evaluation outcome | Controller action |
|---|---|---|---|
| A | On an unsupported storage version, saves an empty version-1 envelope and returns read-failure. | TD-08 fails its preservation assertion. Other ten cases pass by the stated assumption. | Return the finding to Build; one evaluation remains. |
| B | Returns read-failure without saving over unsupported stored data. | All four TD-08 fixtures pass; the other ten cases pass by the stated assumption. | Hand the candidate and evidence to human review. |
The failed observation
For candidate A, the first TD-08 fixture begins with {"version":2,"todos":[]}. The stipulated (illustrative) observations are:
| Observation | Expected | Stipulated for candidate A |
|---|---|---|
| Returned result | ok: false, error: "read-failure" | Matches |
| Stored bytes | {"version":2,"todos":[]} | {"version":1,"todos":[]} |
| Save calls | 0 | 1 |
The result check passes before the preservation check fails. Because that assertion throws, this attempt supplies no result for the remaining three fixtures inside TD-08. It still supplies a decisive counterexample to the required preservation behavior.
Feedback returned to Build
Illustrative feedback: TD-08 failed on the unsupported-version fixture. The constructor returned the expected
read-failure, but it replaced the original version-2 bytes with an empty version-1 envelope and called save once. The obligation requires preserving the original bytes. Repair the failed-read path, retain the contract and protected checks, and submit a new candidate for evaluation. One candidate evaluation remains.
This feedback identifies the behavior to repair and the evidence for the finding. It makes no claim about a failure-mode cause.
The repair and reevaluation
In the illustration, candidate B removes the storage write from the unsupported-data branch. The evaluator runs the protected suite on B, including the other TD-08 fixtures and adjacent storage obligations. A previous pass on A is not reused as evidence for changed candidate B.
The illustrative all-pass result allows the review handoff. If B instead failed and no attempts remained, the controller would stop without qualifying it. A blocking interpretation question would also pause the run for human judgment, even if another attempt remained. Neither the stop nor the pause should discard the candidate or the preceding findings.
What a real record must supply
Replace A and B with exact candidate identities and links to their differences. Bind each result to the contract, check versions, tool configuration and evaluation attempt. Attach the recorded actual and expected observations, execution failures, repair feedback and available timing/token data. Preserve unavailable measurements explicitly.
A report may summarize those records, but the reviewer must be able to follow its links back to them. This illustrative history cannot be used as evidence that the current sample runtime observed this defect or repaired it.
Continue to the human review packet for this story.