The outer, inner, build and learning loops
What the outcome tells us
Terminal evaluation asks whether a completed attempt produced a result that meets the stated criteria. For code, that might mean checking a submitted candidate against a conformance contract. The result can be pass/fail, partial credit or several scores. It tells us about the outcome; that result alone does not explain how the agent arrived there.
A headline success rate across a benchmark summarizes performance on its task set, but can hide differences in repeatability between tasks. Princeton’s reliability study argues for repeated runs and varied conditions instead of relying on one aggregate accuracy score.
The unit being tested matters. A vendor agent combines a model with the vendor’s harness. Results apply to that configuration, with its tools, context, settings and budget. A benchmark using a research or custom harness evaluates that different configuration; its score cannot automatically stand in for the vendor agent’s performance.
Capability asks: can this configured agent perform the task? A verified success demonstrates that it could produce an acceptable result under those conditions. That is useful evidence of capability, but does not establish dependable repetition.
Reliability asks: how dependably does it succeed on this responsibility under the conditions we care about? Repeating a task in fresh trials with the same configuration, starting state, evaluator and budget lets us estimate its success rate. For illustration, 92 successes in 100 trials gives an observed rate of 92% for that case—not a guarantee about the next attempt or a different task.
A hundred or a thousand trials can strengthen that estimate; neither number is a universal requirement. Report the trial count and uncertainty, then test representative cases and relevant variations before generalizing to a capability. Repetition measures dependability on a case; broader coverage tests where that dependability holds.
Both questions can be investigated through terminal outcomes. τ-bench, for example, evaluates final database state and required response content, then uses repeated trials to assess consistent success. Reliability research also examines behavior, resource use and consequences; a success rate is one part of that picture.
These outcomes matter because they provide the evidence on which we decide whether to delegate work or accept a candidate. They do not, by themselves, reveal intermediate errors, recovery or wasted effort. The diagram below shows how we produce and evaluate candidates; the build-loop diagram opens up the construction process.
The loop boundaries
Following Addy Osmani’s Own the Outer Loop, we use inner loop for automated execution and outer loop for human direction and review. In this map, the inner loop spans Build and Evaluate, including controller-driven retries and stopping. The outer loop includes Human Design, delegation to that inner loop and Human Review.
We name two further activities explicitly: the build loop, where an agent, pipeline or graph produces a candidate, and the learning loop, which uses evidence from execution, human review and post-release defects to propose improvements to any phase. Humans approve those process changes. The evaluator sits outside and independent of the build loop, but inside the inner loop. These boundaries define our map; they are not a universal taxonomy.
SDD and loop engineering
Spec-driven development (SDD) and loop engineering (LE) share much of the same workflow. A developer describes the desired end state, then delegates construction to an AI agent that can plan, write code, run checks and revise its work until it believes a candidate is ready. The overlap includes the agent’s iterative work and the process for checking its result. GitHub’s account of SDD, for example, includes implementation plans, tests and feedback as well as the specification itself.
The useful contrast is one of emphasis. SDD makes the specification the organizing artifact; Osmani’s description of loop engineering foregrounds the system that assigns work, checks results, preserves state and decides what happens next. These emphases can coexist in the same implementation. Neither label by itself tells us how independent evaluation is, what ends a run or how much human attention is required. Those are mechanisms we have to examine. The cards below describe our choices for these responsibilities; they do not depend on treating LE as a replacement for SDD.
01 · Before construction
Human Design
Describe what a candidate must do, what existing behavior it must preserve and what it must never permit. Humans own these decisions; AI can help draft obligations, expose ambiguities and propose checks. Preparing and validating those checks belongs to this phase, even when they live in separate files.
Practices to apply
- Make obligations explicit. Give each an identifier, operating conditions, examples and counterexamples. Version the contract so construction and evaluation use the same requirements.
- Connect obligations to evidence. Specify how the evaluator will check each one. Tests, static analysis and formal verification establish different things under different assumptions; identify any judgment that remains with a human.
- Check the required effect. Exercise the relevant interface, prohibited behavior and important state transitions. Challenge the checks with known violations, and confirm that conforming examples can pass.
- Preserve the standard. Keep the approved checks and expected results outside the construction agent’s authority to change.
Remaining limits
A complete-looking contract can omit intent or encode a mistaken assumption. Checks can be weak or wrong, and finite tests generally leave unexamined behavior. EvalPlus exposed incorrect generated programs that passed the original HumanEval tests by expanding the test suite.
Evidence to retain
- The versioned obligations, their rationale, assumptions, exclusions and unresolved questions. Distinguish approved decisions from inferred ones.
- The obligation-to-check mapping, conditions covered and known violations the checks reject, plus any remaining need for human judgment.
Further reading
- Addy Osmani · How to write a good spec for AI agents — goals, boundaries and success criteria.
- GitHub Spec Kit · Specification-Driven Development — specifications, acceptance criteria and deriving tests.
02 · Candidate construction
Build
The agent, pipeline or graph reads context, plans, writes code, uses tools, checks and revises while producing a candidate. The build-loop diagram expands this process. Tool results and checks run during construction provide feedback for its next actions.
Practices to apply
- Supply relevant context and working tools. Preserve the original obligations through intermediate plans.
- Keep attempts traceable. Retain candidate changes, meaningful attempts and observed check results. Submit an identified candidate for external evaluation.
- Keep the handoff explicit. Treat the agent’s completion assessment as a decision to submit work. Qualification still requires the external evaluator and stopping rules.
Handoff to evaluation
For a single-agent build:
- In the ordinary completion path, the agent judges its candidate ready, ends its turn and returns control.
- The controller captures the candidate and calls the external evaluator.
- If evaluation finds unmet obligations and budget remains, the controller returns the findings to the agent for rework.
These are separate decisions: the agent decides to submit; evaluation and stopping rules determine what happens next. Blocked or interrupted attempts may also return control without a completion claim.
Remaining limits
An agent can misuse information it already has. Errors can propagate or recover; a polished completion message establishes neither outcome. Failure as a Process examines how these errors develop through execution.
Confident and Wrong documents repeated submission of patches that fail external tests. Its measured completion signal is patch submission, rather than an explicit verbal claim that every obligation was met.
Evidence to retain
- Exact candidate changes, relevant decisions, tool and check results, failed attempts, corrections and unresolved findings.
- The agent configuration and available timing and token usage, including unsuccessful work, so later comparisons can account for its cost.
Further reading
- Anthropic · Building effective agents — the “Agents” section describes planning, tool use, environmental feedback and termination; Appendix 1 applies the pattern to coding.
03 · External evaluation
Evaluate
One reusable evaluation procedure applies contract-specific checks to an identified candidate and records what the evidence establishes for each obligation. This phase determines whether the candidate is ready for human review, needs repair, has exhausted its retries or needs human judgment before work can continue. The evaluator supplies the findings; controller code carries out the resulting route using those findings and the run’s recorded limits.
Practices to apply
- Protect the evaluation. The evaluator is outside the build loop and inside the inner loop. Keep authoritative checks, expected results, verdicts and stopping rules outside the producing agent, pipeline or graph’s authority to change.
- Evaluate the submitted candidate. Bind results to its exact contents, the contract and check versions, tools and operating conditions. Re-evaluate when the candidate changes.
- Account for every obligation. Distinguish satisfied, not satisfied and unverified. Missing, failed or invalid required evidence cannot become a pass. Explicitly identify any judgments reserved for human review.
- Enforce limits in the controller. Bound attempts and resources, define how to handle unavailable checks or lack of progress, and preserve unsuccessful outcomes.
Four possible outcomes
- Ready for human review. The candidate meets the declared prerequisites for review. Pass it to the human with the supporting evidence.
- Return for repair. The candidate does not satisfy identified obligations and another attempt is permitted. Return it with a note identifying those obligations, the observed failures and the evidence needed to assess a repair.
- Retries exhausted. The candidate still needs work and no retries remain. Stop the run and preserve the candidate, findings and attempt history without qualifying it.
- Human judgment required. A blocking question or condition cannot be resolved automatically. Pause execution and raise the specific matter to a human with the relevant evidence and decision needed.
The last route is distinct from the planned review of a candidate that is ready. Record other resource limits or unavailable checks explicitly as well; a pause is not a successful evaluation.
Remaining limits
Protecting the evaluator’s authority does not make its checks complete or correct. Weak assertions, shared mistaken assumptions, tool defects or inaccurate AI judgments can still admit flawed code. Deterministic stopping faithfully enforces its inputs; it cannot strengthen the evidence those inputs contain.
Evidence to retain
- Candidate and contract identities; check versions, execution conditions, actual outputs and per-obligation findings, including unavailable evidence.
- The stopping rule applied, its inputs and decision, plus evaluation history and costs. A final pass should not erase earlier failures.
Further reading
- Sydney Runkle · The Art of Loop Engineering — the verification loop checks an agent’s output and returns findings for rework.
- Yoko Li · Knowing When to Stop — completion criteria, verifier limitations and the cost of continued iteration.
04 · Human judgment
Human Review
Bring the candidate, its obligations, supporting evidence and remaining uncertainty together so a human can make an informed decision. AI can assemble this material even when it cannot verify an obligation itself. The aim is to preserve human time and attention by reducing the work of finding and connecting the evidence.
Practices to apply
- Organize the evidence around obligations. Show the expected behavior, relevant code changes, check results and limitations together. Link directly to the code, contract, design decisions and underlying records.
- Direct attention to consequential gaps. Identify assumptions, unverified behavior and possible omissions from the contract. Highlight integration, security, concurrency or user-experience questions where the available checks provide weak evidence.
- Make uncertainty inspectable. Separate observed results from AI interpretations. Explain what each check establishes, under which conditions, and what remains unknown. Keep contrary findings visible.
- Ask for specific decisions. State which behavior, assumption or tradeoff needs judgment and supply the relevant evidence. The reviewer can accept, request changes or require further investigation. Return implementation changes to Build and re-evaluate them; revisit Human Design when the contract or its checks need revision.
Remaining limits
An AI summary can omit a defect, misstate a result or direct attention away from the important issue. Keep the underlying evidence accessible and let reviewers widen their inspection. Human review can miss errors too. A polished evidence package alone does not establish that review is faster or more effective; those outcomes need measurement.
Evidence to retain
- The exact candidate and evidence package reviewed, including the scope of automated checks and the questions requiring human judgment.
- The reviewer’s decision, rationale, additional checks, requested changes and unresolved limitations. Keep evaluator qualification and human acceptance distinct.
Further reading
- Addy Osmani · Own the Outer Loop — evidence, human judgment and ownership at the boundary of automated work.
- Google Engineering Practices · What to look for in a code review — design, behavior, test quality and the surrounding system context.