One process, four methods

These four approaches share a basic structure: define the task, let the system attempt it, and assess the result. Each emphasizes different practices, but those practices can be combined in the same system.

1 · Specify

Define what the result must do, preserve and never permit.

2 · Build

An agent, pipeline or graph attempts the task and submits a candidate.

In this view, construction remains a closed box.

3 · Evaluate

Assess the candidate against the stated criteria.

Result: pass, fail, a score or an unresolved assessment.

Optional routes after evaluation

  • Qualified → advanceHuman review or the next designated stage.
  • Rejected → reworkReturn to Build ↶ with findings, if retries remain.
  • Retries exhausted → stopStop without a qualified candidate.
  • Judgment needed → escalatePause and raise the question to a human.

Any of the four approaches can use these routes. The surrounding workflow decides what follows evaluation; not every task requires every route.

Agent-based model evaluationAsks what a configured agent can do and how dependably, from the outcomes of repeated attempts.Explore practices and citations
End result specification

A task statement, a starting repository and the checks that define success

Autonomous build

The configured agent attempts the task

  • For repeated trials of one task, hold configuration, starting state and budget fixed7
Evaluation

The checks score the final state

  • Repeat in fresh trials; report the success rate and its uncertainty7
  • Protect benchmark evaluation material from contamination8
Prompt engineeringEmphasizes the instruction the model receives: clear wording, examples and context, improved against success criteria.Explore practices and citations
End result specification

A prompt: the task, success criteria and relevant context

  • Success criteria and empirical tests defined first1
  • Clear instructions, examples, structure and role1
Autonomous build

The model, or an agent, responds

  • Thinking and prompt chaining1
Evaluation

The output is judged against the success criteria

  • Test empirically against the criteria; revise the prompt1
Spec-driven development (SDD)Makes the written specification the organizing artifact; code is generated from it and checked against it.Explore practices and citations
End result specification

A specification with acceptance criteria

  • Specification as the source of truth: requirements, plan, tasks2
  • Ambiguities flagged and clarified; checklists test the specification2
  • A constitution of standing principles and gates2
Autonomous build

An agent implements from the plan

  • Implement task by task from the plan2
Evaluation

The result is checked against the specification and its tests

  • Tests written and approved before implementation2
  • The human's role is to verify, not only to steer3
Loop engineeringEmphasizes the system around the agent: who starts work, who checks it, what ends a run and what carries over.Explore practices and citations
End result specification

A goal and its acceptance criteria

  • A goal and verifiable stop condition written up front4, 5
Autonomous build

An agent, pipeline or graph builds

  • Automated handoffs between agents and stages4
  • State kept outside the conversation, across attempts4
  • Triggers or schedules start the runs4, 5
Evaluation

A verifier checks the result

  • An evaluator separate from the builder4
  • Turn and retry limits; escalate to a human5, 6
  • Findings returned for another attempt6
This view tells us what task the system received and how its result was judged. It does not tell us which errors arose during construction, whether the system recovered, or what that recovery cost. We will open that box next.
Sources and scope of the comparison

Selected guidance. This is an editorial comparison of emphasis, not a survey of adoption or a priority claim.

A prompted agent can check and repair its work during an attempt, and one benchmark attempt can contain an entire construction loop (SWE-agent). Spec Kit supplies the SDD examples shown here; they are not requirements of every SDD workflow.

  1. Anthropic, Prompt engineering overview: success criteria and empirical tests first; clarity, examples, structure, role prompting, thinking and prompt chaining.
  2. GitHub Spec Kit, Specification-Driven Development: specification, plan and tasks; clarification markers and checklists; constitution and gates; test-first imperative.
  3. GitHub Blog, Spec-driven development with AI (2 Sep 2025): the developer's role is to verify.
  4. Addy Osmani, Loop Engineering (7 Jun 2026): handoffs, a separate checker, a condition you wrote, state files and automations.
  5. Macedo, Stop Hand-Holding Your Coding Agent (28 Jun 2026, arXiv 2607.00038): a loop specification with trigger, goal, verification, stopping rule and memory; human escalation.
  6. Anthropic, Loop engineering: Getting started with loops (30 Jun 2026): evaluator-checked goals and explicit turn caps. Vendor material.
  7. Princeton reliability study and τ-bench: repeated trials and consistent success under comparable conditions.
  8. Wang et al., SDLC survey of code-LLM benchmarks (v3, 6 Mar 2026): calls for effective anti-contamination mechanisms.