# The agentic code factory: from request to releasable change

By Vibe Haus · Published and source-reviewed 2026-10-10

Source: https://vibehaus.team/blog/agentic-code-factory

A practical design for an agentic code factory: task contracts, isolated execution, evidence, review gates, bounded retries, and release ownership.

Original Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed.

> A useful code factory produces changes a team can trust. Its unit of work is a releasable change with evidence, an owner, and a recovery path.

## What the factory actually produces

An agent can turn a request into a patch quickly enough to make coding look like the whole job. The surrounding decisions remain: what should change, which behavior must survive, whether the patch works in its real environment, and who accepts the consequences of shipping it. Automating the patch without those decisions moves unfinished work into someone else’s queue.

Our proposed code factory is a delivery system with explicit states. A request enters as an unresolved problem. It leaves as an accepted change, a rejected experiment, or a documented decision to stop. Each transition requires an artifact another person or tool can inspect. “The agent says it is done” is not a transition condition.

Anthropic’s 2024 engineering guidance distinguishes predefined workflows from agents that choose their own next steps. That distinction is useful here: let an agent explore an implementation while ordinary code enforces budgets, permissions, and release gates. The source now notes that its tooling landscape has evolved; this article uses the architectural distinction, not its older tool catalog.

### Workflow: A proposed delivery path. A failed gate never silently becomes permission to ship.

1. Frame: Agree on intent and acceptance evidence.
2. Build: Work within an isolated, bounded task.
3. Verify: Run checks and assemble revision-bound evidence.

Decision: Does this revision meet the contract?
- Pass → independent review, then authorized release.
- Fail → diagnose and retry within budget; otherwise return to the owner.

Sources: [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents)

## Make the request testable before making it executable

Sean Grove’s The New Code argues for preserving structured intent instead of treating code as the only durable record of a change. Our implementation of that idea is a small task contract: the user outcome, constraints, acceptance examples, excluded work, and release owner. It should be short enough to review alongside the diff and specific enough to reject a plausible but wrong implementation.

Consider adding CSV export to a customer dashboard. “Add export” leaves permissions, date filtering, column order, large datasets, and spreadsheet formula handling undefined. A useful contract names the authorized user, the records that may appear, the expected output, and the behavior when the export is too large. It also records whether synchronous export is acceptable or a background job is required.

Do not demand a perfect specification for uncertain work. Mark unknowns explicitly. A feasibility task can finish with a measured limitation and a recommendation instead of production code. The contract must distinguish an experiment from a commitment to deliver a feature; otherwise a prototype tends to inherit production expectations without production checks.

```yaml
# Illustrative task contract; adapt to the actual product.
outcome: An account admin can export the current filtered view.
acceptance:
  - Export includes only records the caller may read.
  - Column order and timezone match the agreed sample.
  - A second account cannot export the first account's records.
unknowns:
  - Maximum supported export size; measure before choosing sync or async.
excluded:
  - Scheduled reports and third-party delivery.
evidence:
  - Permission regression test and representative output sample.
release_owner: Accountable engineer assigned before release.
stop_when:
  - Required data access or product decisions exceed task authority.
```

Sources: [The New Code](https://ai.engineer/talks/8rABwKRsec4-specifications-are-the-new-code)

## Give each worker a bounded place to work

Use a separate checkout or worktree for each concurrent code task. Record the base commit and dependencies at creation. A worktree prevents accidental file collisions; it does not isolate processes, credentials, or the network. If untrusted code must run, pair the checkout with an execution boundary appropriate to the risk, such as a constrained container or sandbox.

The worker needs the smallest useful set of capabilities: read the relevant repository, edit the intended area, and run the approved checks. A documentation task usually has no reason to receive production credentials. A migration task may need representative schema information without access to customer rows. Keep authority in the runtime and account configuration rather than relying on prose to enforce it.

Parallel work is justified when tasks have independent acceptance conditions. Two workers editing the same central schema can create an integration bottleneck even when each patch is locally correct. Split along stable interfaces, publish those interfaces first, or serialize the coupled changes. Account for integration work in the plan instead of treating it as a free final step.

- Record base revision, permitted tools, task owner, and maximum execution budget.
- Give each task a visible state: ready, running, awaiting decision, verifying, reviewing, or accepted.
- Stop and return a concrete question when missing authority or a product decision prevents progress.

## Treat evidence as a build artifact

The evidence packet should travel with the candidate change. Include the base and head commits, a concise description of the behavior change, checks actually executed, their results, and important checks that were not run. A browser screenshot can support a layout claim. It cannot establish authorization correctness, database safety, or keyboard behavior by itself.

For the export example, useful evidence includes a permitted export, a cross-account denial, a sample output checked against the contract, and a measurement near the agreed size boundary. If the agent only ran unit tests around a helper function, say so. Do not translate “helper test passed” into “export flow verified.”

Bind the packet to the exact candidate revision. A later fix, merge, dependency change, or generated-file update can invalidate part of the evidence. Define which checks must rerun when those inputs change. Store durable logs or reports where the reviewer can inspect them; a prose claim without its underlying result is a weaker artifact.

- Behavior: what changed and what remained an explicit constraint.
- Provenance: base revision, candidate revision, toolchain, and relevant environment.
- Verification: commands, results, representative examples, and omissions.
- Decision: unresolved risks, assigned owner, and proposed release or stop condition.

## Control work in progress before adding more agents

A factory has queues. More implementation workers can increase the arrival rate at review while the review team’s capacity stays fixed. The result is older branches, repeated rebases, and more context switching. Track elapsed time from request to acceptance alongside active work time; a faster implementation stage can coexist with slower delivery.

Set a work-in-progress limit at the narrowest stage. For example, a team might initially allow only two review-ready changes per reviewer. That number is an illustrative starting policy, not a universal capacity rule. Adjust it using the observed age of the review queue and the complexity of the changes, rather than maximizing the count of simultaneously running agents.

Bound repair loops as well. A failing check should produce a diagnosis before another attempt. Repeated changes that alternate between two failures indicate missing understanding, an unstable test, or a bad task boundary. Stop after the agreed attempt or cost budget, preserve the evidence, and ask the owner to change the task or investigate. A silent infinite loop is an operational failure even if it eventually produces a patch.

## Review and release are separate transitions

An independent reviewer should reconstruct the important behavior from the code, requirements, and evidence. Independence means a distinct verification path; merely using a second model does not guarantee it. The reviewer can share the specification with the author while choosing different boundary cases and inspecting the underlying results directly.

Acceptance does not automatically grant deployment authority. The release step should use the project’s existing authorization rules, target the reviewed revision, and name the person or automation responsible for monitoring. If the integration branch moved, reconcile that change before reusing earlier approvals. Database or external API changes may need compatibility sequencing rather than a single all-at-once deployment.

Recovery must be concrete. A frontend rollback may be simple; a data migration may require a forward repair. Record the signal that would trigger intervention, the owner who watches it, and the viable recovery action. For the export feature, elevated authorization failures or incorrect row exposure matter more than whether the deployment process returned a success code.

## Start with one repeatable task type

Choose a task family with a clear outcome and observable checks: a small UI change, a well-scoped bug, or an internal automation. Run the proposed process with one worker and a human reviewer before adding orchestration. Capture where the task stalls and which artifacts the reviewer actually uses. Remove ceremonial paperwork; strengthen missing evidence.

Expand only after the team can explain a failed run. Was the request ambiguous, context stale, implementation incorrect, verification incomplete, or release handling weak? These require different repairs. Buying a stronger model may help some implementation failures while leaving all the other failure classes untouched.

The process needs five things: contracts, isolated execution, revision-bound evidence, bounded queues, and accountable decisions. Our downloadable task-contract skill is a starting template for the first stage. The companion code-review and evaluation articles explain how to challenge the output and decide whether the overall system is worth expanding.

## Sources and further reading

Sources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations.

- [The New Code](https://ai.engineer/talks/8rABwKRsec4-specifications-are-the-new-code): Sean Grove · AI Engineer Worlds Fair 2025. Conference talk.

- [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents): Anthropic · 2024, with subsequent tooling notice. Engineering guidance.

## Related services

- [ai native engineering](https://vibehaus.team/ai-native-engineering)

- [managed engineering](https://vibehaus.team/managed-engineering)

- [how we work](https://vibehaus.team/how-we-work)

## Reuse

You may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms.

Downloads and tools: https://vibehaus.team/data