One process, four methods
These four approaches share a basic structure: define the task, let the system attempt it, and assess the result. Each emphasizes different practices, but those practices can be combined in the same system.
- The shared process
- Practices emphasized by the cited guidance—not exclusive to this approach.
Define what the result must do, preserve and never permit.
An agent, pipeline or graph attempts the task and submits a candidate.
In this view, construction remains a closed box.
Assess the candidate against the stated criteria.
Result: pass, fail, a score or an unresolved assessment.
Optional routes after evaluation
- Qualified → advanceHuman review or the next designated stage.
- Rejected → reworkReturn to Build ↶ with findings, if retries remain.
- Retries exhausted → stopStop without a qualified candidate.
- Judgment needed → escalatePause and raise the question to a human.
Any of the four approaches can use these routes. The surrounding workflow decides what follows evaluation; not every task requires every route.
Agent-based model evaluationAsks what a configured agent can do and how dependably, from the outcomes of repeated attempts.Explore practices and citations
A task statement, a starting repository and the checks that define success
The configured agent attempts the task
- For repeated trials of one task, hold configuration, starting state and budget fixed7
Prompt engineeringEmphasizes the instruction the model receives: clear wording, examples and context, improved against success criteria.Explore practices and citations
A prompt: the task, success criteria and relevant context
The output is judged against the success criteria
- Test empirically against the criteria; revise the prompt1
Spec-driven development (SDD)Makes the written specification the organizing artifact; code is generated from it and checked against it.Explore practices and citations
A specification with acceptance criteria
Loop engineeringEmphasizes the system around the agent: who starts work, who checks it, what ends a run and what carries over.Explore practices and citations
A goal and its acceptance criteria
An agent, pipeline or graph builds
Sources and scope of the comparison
Selected guidance. This is an editorial comparison of emphasis, not a survey of adoption or a priority claim.
A prompted agent can check and repair its work during an attempt, and one benchmark attempt can contain an entire construction loop (SWE-agent). Spec Kit supplies the SDD examples shown here; they are not requirements of every SDD workflow.
- Anthropic, Prompt engineering overview: success criteria and empirical tests first; clarity, examples, structure, role prompting, thinking and prompt chaining.
- GitHub Spec Kit, Specification-Driven Development: specification, plan and tasks; clarification markers and checklists; constitution and gates; test-first imperative.
- GitHub Blog, Spec-driven development with AI (2 Sep 2025): the developer's role is to verify.
- Addy Osmani, Loop Engineering (7 Jun 2026): handoffs, a separate checker, a condition you wrote, state files and automations.
- Macedo, Stop Hand-Holding Your Coding Agent (28 Jun 2026, arXiv 2607.00038): a loop specification with trigger, goal, verification, stopping rule and memory; human escalation.
- Anthropic, Loop engineering: Getting started with loops (30 Jun 2026): evaluator-checked goals and explicit turn caps. Vendor material.
- Princeton reliability study and τ-bench: repeated trials and consistent success under comparable conditions.
- Wang et al., SDLC survey of code-LLM benchmarks (v3, 6 Mar 2026): calls for effective anti-contamination mechanisms.