{"schemaVersion":"1.0","updated":"2026-10-10","publisher":"Vibe Haus","method":"Original Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed.","url":"https://vibehaus.team/blog","mcp":"https://vibehaus.team/mcp","articles":[{"slug":"agentic-code-factory","title":"The agentic code factory: from request to releasable change","description":"A practical design for an agentic code factory: task contracts, isolated execution, evidence, review gates, bounded retries, and release ownership.","topic":"Delivery systems","date":"2026-10-10","thesis":"A useful code factory produces changes a team can trust. Its unit of work is a releasable change with evidence, an owner, and a recovery path.","sections":[{"id":"factory","title":"What the factory actually produces","paragraphs":["An agent can turn a request into a patch quickly enough to make coding look like the whole job. The surrounding decisions remain: what should change, which behavior must survive, whether the patch works in its real environment, and who accepts the consequences of shipping it. Automating the patch without those decisions moves unfinished work into someone else’s queue.","Our proposed code factory is a delivery system with explicit states. A request enters as an unresolved problem. It leaves as an accepted change, a rejected experiment, or a documented decision to stop. Each transition requires an artifact another person or tool can inspect. “The agent says it is done” is not a transition condition.","Anthropic’s 2024 engineering guidance distinguishes predefined workflows from agents that choose their own next steps. That distinction is useful here: let an agent explore an implementation while ordinary code enforces budgets, permissions, and release gates. The source now notes that its tooling landscape has evolved; this article uses the architectural distinction, not its older tool catalog."],"flow":{"caption":"A proposed delivery path. A failed gate never silently becomes permission to ship.","steps":[{"title":"Frame","detail":"Agree on intent and acceptance evidence."},{"title":"Build","detail":"Work within an isolated, bounded task."},{"title":"Verify","detail":"Run checks and assemble revision-bound evidence."}],"gate":"Does this revision meet the contract?","pass":"Pass → independent review, then authorized release.","fail":"Fail → diagnose and retry within budget; otherwise return to the owner."},"sources":["anthropic"]},{"id":"contract","title":"Make the request testable before making it executable","paragraphs":["Sean Grove’s The New Code argues for preserving structured intent instead of treating code as the only durable record of a change. Our implementation of that idea is a small task contract: the user outcome, constraints, acceptance examples, excluded work, and release owner. It should be short enough to review alongside the diff and specific enough to reject a plausible but wrong implementation.","Consider adding CSV export to a customer dashboard. “Add export” leaves permissions, date filtering, column order, large datasets, and spreadsheet formula handling undefined. A useful contract names the authorized user, the records that may appear, the expected output, and the behavior when the export is too large. It also records whether synchronous export is acceptable or a background job is required.","Do not demand a perfect specification for uncertain work. Mark unknowns explicitly. A feasibility task can finish with a measured limitation and a recommendation instead of production code. The contract must distinguish an experiment from a commitment to deliver a feature; otherwise a prototype tends to inherit production expectations without production checks."],"code":{"language":"yaml","text":"# Illustrative task contract; adapt to the actual product.\noutcome: An account admin can export the current filtered view.\nacceptance:\n  - Export includes only records the caller may read.\n  - Column order and timezone match the agreed sample.\n  - A second account cannot export the first account's records.\nunknowns:\n  - Maximum supported export size; measure before choosing sync or async.\nexcluded:\n  - Scheduled reports and third-party delivery.\nevidence:\n  - Permission regression test and representative output sample.\nrelease_owner: Accountable engineer assigned before release.\nstop_when:\n  - Required data access or product decisions exceed task authority."},"sources":["grove"]},{"id":"execution","title":"Give each worker a bounded place to work","paragraphs":["Use a separate checkout or worktree for each concurrent code task. Record the base commit and dependencies at creation. A worktree prevents accidental file collisions; it does not isolate processes, credentials, or the network. If untrusted code must run, pair the checkout with an execution boundary appropriate to the risk, such as a constrained container or sandbox.","The worker needs the smallest useful set of capabilities: read the relevant repository, edit the intended area, and run the approved checks. A documentation task usually has no reason to receive production credentials. A migration task may need representative schema information without access to customer rows. Keep authority in the runtime and account configuration rather than relying on prose to enforce it.","Parallel work is justified when tasks have independent acceptance conditions. Two workers editing the same central schema can create an integration bottleneck even when each patch is locally correct. Split along stable interfaces, publish those interfaces first, or serialize the coupled changes. Account for integration work in the plan instead of treating it as a free final step."],"bullets":["Record base revision, permitted tools, task owner, and maximum execution budget.","Give each task a visible state: ready, running, awaiting decision, verifying, reviewing, or accepted.","Stop and return a concrete question when missing authority or a product decision prevents progress."]},{"id":"evidence","title":"Treat evidence as a build artifact","paragraphs":["The evidence packet should travel with the candidate change. Include the base and head commits, a concise description of the behavior change, checks actually executed, their results, and important checks that were not run. A browser screenshot can support a layout claim. It cannot establish authorization correctness, database safety, or keyboard behavior by itself.","For the export example, useful evidence includes a permitted export, a cross-account denial, a sample output checked against the contract, and a measurement near the agreed size boundary. If the agent only ran unit tests around a helper function, say so. Do not translate “helper test passed” into “export flow verified.”","Bind the packet to the exact candidate revision. A later fix, merge, dependency change, or generated-file update can invalidate part of the evidence. Define which checks must rerun when those inputs change. Store durable logs or reports where the reviewer can inspect them; a prose claim without its underlying result is a weaker artifact."],"bullets":["Behavior: what changed and what remained an explicit constraint.","Provenance: base revision, candidate revision, toolchain, and relevant environment.","Verification: commands, results, representative examples, and omissions.","Decision: unresolved risks, assigned owner, and proposed release or stop condition."]},{"id":"queues","title":"Control work in progress before adding more agents","paragraphs":["A factory has queues. More implementation workers can increase the arrival rate at review while the review team’s capacity stays fixed. The result is older branches, repeated rebases, and more context switching. Track elapsed time from request to acceptance alongside active work time; a faster implementation stage can coexist with slower delivery.","Set a work-in-progress limit at the narrowest stage. For example, a team might initially allow only two review-ready changes per reviewer. That number is an illustrative starting policy, not a universal capacity rule. Adjust it using the observed age of the review queue and the complexity of the changes, rather than maximizing the count of simultaneously running agents.","Bound repair loops as well. A failing check should produce a diagnosis before another attempt. Repeated changes that alternate between two failures indicate missing understanding, an unstable test, or a bad task boundary. Stop after the agreed attempt or cost budget, preserve the evidence, and ask the owner to change the task or investigate. A silent infinite loop is an operational failure even if it eventually produces a patch."]},{"id":"release","title":"Review and release are separate transitions","paragraphs":["An independent reviewer should reconstruct the important behavior from the code, requirements, and evidence. Independence means a distinct verification path; merely using a second model does not guarantee it. The reviewer can share the specification with the author while choosing different boundary cases and inspecting the underlying results directly.","Acceptance does not automatically grant deployment authority. The release step should use the project’s existing authorization rules, target the reviewed revision, and name the person or automation responsible for monitoring. If the integration branch moved, reconcile that change before reusing earlier approvals. Database or external API changes may need compatibility sequencing rather than a single all-at-once deployment.","Recovery must be concrete. A frontend rollback may be simple; a data migration may require a forward repair. Record the signal that would trigger intervention, the owner who watches it, and the viable recovery action. For the export feature, elevated authorization failures or incorrect row exposure matter more than whether the deployment process returned a success code."]},{"id":"adoption","title":"Start with one repeatable task type","paragraphs":["Choose a task family with a clear outcome and observable checks: a small UI change, a well-scoped bug, or an internal automation. Run the proposed process with one worker and a human reviewer before adding orchestration. Capture where the task stalls and which artifacts the reviewer actually uses. Remove ceremonial paperwork; strengthen missing evidence.","Expand only after the team can explain a failed run. Was the request ambiguous, context stale, implementation incorrect, verification incomplete, or release handling weak? These require different repairs. Buying a stronger model may help some implementation failures while leaving all the other failure classes untouched.","The process needs five things: contracts, isolated execution, revision-bound evidence, bounded queues, and accountable decisions. Our downloadable task-contract skill is a starting template for the first stage. The companion code-review and evaluation articles explain how to challenge the output and decide whether the overall system is worth expanding."]}],"sources":[{"id":"grove","title":"The New Code","author":"Sean Grove · AI Engineer Worlds Fair 2025","url":"https://ai.engineer/talks/8rABwKRsec4-specifications-are-the-new-code","kind":"Conference talk"},{"id":"anthropic","title":"Building effective agents","author":"Anthropic · 2024, with subsequent tooling notice","url":"https://www.anthropic.com/engineering/building-effective-agents","kind":"Engineering guidance"}],"url":"https://vibehaus.team/blog/agentic-code-factory","markdownUrl":"https://vibehaus.team/blog/agentic-code-factory/article.md","markdown":"# The agentic code factory: from request to releasable change\n\nBy Vibe Haus · Published and source-reviewed 2026-10-10\n\nSource: https://vibehaus.team/blog/agentic-code-factory\n\nA practical design for an agentic code factory: task contracts, isolated execution, evidence, review gates, bounded retries, and release ownership.\n\nOriginal Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed.\n\n> A useful code factory produces changes a team can trust. Its unit of work is a releasable change with evidence, an owner, and a recovery path.\n\n## What the factory actually produces\n\nAn agent can turn a request into a patch quickly enough to make coding look like the whole job. The surrounding decisions remain: what should change, which behavior must survive, whether the patch works in its real environment, and who accepts the consequences of shipping it. Automating the patch without those decisions moves unfinished work into someone else’s queue.\n\nOur proposed code factory is a delivery system with explicit states. A request enters as an unresolved problem. It leaves as an accepted change, a rejected experiment, or a documented decision to stop. Each transition requires an artifact another person or tool can inspect. “The agent says it is done” is not a transition condition.\n\nAnthropic’s 2024 engineering guidance distinguishes predefined workflows from agents that choose their own next steps. That distinction is useful here: let an agent explore an implementation while ordinary code enforces budgets, permissions, and release gates. The source now notes that its tooling landscape has evolved; this article uses the architectural distinction, not its older tool catalog.\n\n### Workflow: A proposed delivery path. A failed gate never silently becomes permission to ship.\n\n1. Frame: Agree on intent and acceptance evidence.\n2. Build: Work within an isolated, bounded task.\n3. Verify: Run checks and assemble revision-bound evidence.\n\nDecision: Does this revision meet the contract?\n- Pass → independent review, then authorized release.\n- Fail → diagnose and retry within budget; otherwise return to the owner.\n\nSources: [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents)\n\n## Make the request testable before making it executable\n\nSean Grove’s The New Code argues for preserving structured intent instead of treating code as the only durable record of a change. Our implementation of that idea is a small task contract: the user outcome, constraints, acceptance examples, excluded work, and release owner. It should be short enough to review alongside the diff and specific enough to reject a plausible but wrong implementation.\n\nConsider adding CSV export to a customer dashboard. “Add export” leaves permissions, date filtering, column order, large datasets, and spreadsheet formula handling undefined. A useful contract names the authorized user, the records that may appear, the expected output, and the behavior when the export is too large. It also records whether synchronous export is acceptable or a background job is required.\n\nDo not demand a perfect specification for uncertain work. Mark unknowns explicitly. A feasibility task can finish with a measured limitation and a recommendation instead of production code. The contract must distinguish an experiment from a commitment to deliver a feature; otherwise a prototype tends to inherit production expectations without production checks.\n\n```yaml\n# Illustrative task contract; adapt to the actual product.\noutcome: An account admin can export the current filtered view.\nacceptance:\n  - Export includes only records the caller may read.\n  - Column order and timezone match the agreed sample.\n  - A second account cannot export the first account's records.\nunknowns:\n  - Maximum supported export size; measure before choosing sync or async.\nexcluded:\n  - Scheduled reports and third-party delivery.\nevidence:\n  - Permission regression test and representative output sample.\nrelease_owner: Accountable engineer assigned before release.\nstop_when:\n  - Required data access or product decisions exceed task authority.\n```\n\nSources: [The New Code](https://ai.engineer/talks/8rABwKRsec4-specifications-are-the-new-code)\n\n## Give each worker a bounded place to work\n\nUse a separate checkout or worktree for each concurrent code task. Record the base commit and dependencies at creation. A worktree prevents accidental file collisions; it does not isolate processes, credentials, or the network. If untrusted code must run, pair the checkout with an execution boundary appropriate to the risk, such as a constrained container or sandbox.\n\nThe worker needs the smallest useful set of capabilities: read the relevant repository, edit the intended area, and run the approved checks. A documentation task usually has no reason to receive production credentials. A migration task may need representative schema information without access to customer rows. Keep authority in the runtime and account configuration rather than relying on prose to enforce it.\n\nParallel work is justified when tasks have independent acceptance conditions. Two workers editing the same central schema can create an integration bottleneck even when each patch is locally correct. Split along stable interfaces, publish those interfaces first, or serialize the coupled changes. Account for integration work in the plan instead of treating it as a free final step.\n\n- Record base revision, permitted tools, task owner, and maximum execution budget.\n- Give each task a visible state: ready, running, awaiting decision, verifying, reviewing, or accepted.\n- Stop and return a concrete question when missing authority or a product decision prevents progress.\n\n## Treat evidence as a build artifact\n\nThe evidence packet should travel with the candidate change. Include the base and head commits, a concise description of the behavior change, checks actually executed, their results, and important checks that were not run. A browser screenshot can support a layout claim. It cannot establish authorization correctness, database safety, or keyboard behavior by itself.\n\nFor the export example, useful evidence includes a permitted export, a cross-account denial, a sample output checked against the contract, and a measurement near the agreed size boundary. If the agent only ran unit tests around a helper function, say so. Do not translate “helper test passed” into “export flow verified.”\n\nBind the packet to the exact candidate revision. A later fix, merge, dependency change, or generated-file update can invalidate part of the evidence. Define which checks must rerun when those inputs change. Store durable logs or reports where the reviewer can inspect them; a prose claim without its underlying result is a weaker artifact.\n\n- Behavior: what changed and what remained an explicit constraint.\n- Provenance: base revision, candidate revision, toolchain, and relevant environment.\n- Verification: commands, results, representative examples, and omissions.\n- Decision: unresolved risks, assigned owner, and proposed release or stop condition.\n\n## Control work in progress before adding more agents\n\nA factory has queues. More implementation workers can increase the arrival rate at review while the review team’s capacity stays fixed. The result is older branches, repeated rebases, and more context switching. Track elapsed time from request to acceptance alongside active work time; a faster implementation stage can coexist with slower delivery.\n\nSet a work-in-progress limit at the narrowest stage. For example, a team might initially allow only two review-ready changes per reviewer. That number is an illustrative starting policy, not a universal capacity rule. Adjust it using the observed age of the review queue and the complexity of the changes, rather than maximizing the count of simultaneously running agents.\n\nBound repair loops as well. A failing check should produce a diagnosis before another attempt. Repeated changes that alternate between two failures indicate missing understanding, an unstable test, or a bad task boundary. Stop after the agreed attempt or cost budget, preserve the evidence, and ask the owner to change the task or investigate. A silent infinite loop is an operational failure even if it eventually produces a patch.\n\n## Review and release are separate transitions\n\nAn independent reviewer should reconstruct the important behavior from the code, requirements, and evidence. Independence means a distinct verification path; merely using a second model does not guarantee it. The reviewer can share the specification with the author while choosing different boundary cases and inspecting the underlying results directly.\n\nAcceptance does not automatically grant deployment authority. The release step should use the project’s existing authorization rules, target the reviewed revision, and name the person or automation responsible for monitoring. If the integration branch moved, reconcile that change before reusing earlier approvals. Database or external API changes may need compatibility sequencing rather than a single all-at-once deployment.\n\nRecovery must be concrete. A frontend rollback may be simple; a data migration may require a forward repair. Record the signal that would trigger intervention, the owner who watches it, and the viable recovery action. For the export feature, elevated authorization failures or incorrect row exposure matter more than whether the deployment process returned a success code.\n\n## Start with one repeatable task type\n\nChoose a task family with a clear outcome and observable checks: a small UI change, a well-scoped bug, or an internal automation. Run the proposed process with one worker and a human reviewer before adding orchestration. Capture where the task stalls and which artifacts the reviewer actually uses. Remove ceremonial paperwork; strengthen missing evidence.\n\nExpand only after the team can explain a failed run. Was the request ambiguous, context stale, implementation incorrect, verification incomplete, or release handling weak? These require different repairs. Buying a stronger model may help some implementation failures while leaving all the other failure classes untouched.\n\nThe process needs five things: contracts, isolated execution, revision-bound evidence, bounded queues, and accountable decisions. Our downloadable task-contract skill is a starting template for the first stage. The companion code-review and evaluation articles explain how to challenge the output and decide whether the overall system is worth expanding.\n\n## Sources and further reading\n\nSources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations.\n\n- [The New Code](https://ai.engineer/talks/8rABwKRsec4-specifications-are-the-new-code): Sean Grove · AI Engineer Worlds Fair 2025. Conference talk.\n\n- [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents): Anthropic · 2024, with subsequent tooling notice. Engineering guidance.\n\n## Related services\n\n- [ai native engineering](https://vibehaus.team/ai-native-engineering)\n\n- [managed engineering](https://vibehaus.team/managed-engineering)\n\n- [how we work](https://vibehaus.team/how-we-work)\n\n## Reuse\n\nYou may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms.\n\nDownloads and tools: https://vibehaus.team/data"},{"slug":"code-review-for-agents","title":"Code review after agents: verify the change, not the explanation","description":"An evidence-first review workflow for agent-written code, covering risk routing, independent tests, actionable findings, and exact-revision release gates.","topic":"Quality & review","date":"2026-10-10","thesis":"Agent-written code deserves the same engineering standard as any other change. Faster generation makes review capacity, evidence quality, and risk selection more important.","sections":[{"id":"review-object","title":"Review three things: intent, implementation, evidence","paragraphs":["A convincing pull request description can make a change feel understood before the reviewer has read the code. Start by separating three things. Intent says what the user needs and what must remain true. Implementation is the actual diff and its surrounding behavior. Evidence shows what was checked against that intent. Agreement between all three is stronger than a polished explanation of any one.","Google’s engineering review guidance covers design, functionality, complexity, tests, naming, comments, style, and documentation. Agent involvement does not remove those concerns. Our proposed workflow changes the order: establish the risk and the behavioral claim first, then spend review attention where a plausible patch could still be wrong.","For example, an agent may add a permission check to a new export route and describe the route as secure. The reviewer still needs to determine whose account the query uses, whether the caller can choose another account identifier, and whether the underlying data access enforces the same boundary. A correct-looking guard is not evidence that the entire path respects it."],"flow":{"caption":"Review attention follows the possible consequence of the change.","steps":[{"title":"Route risk","detail":"Identify data, permissions, money, availability, and user-visible behavior."},{"title":"Inspect","detail":"Read the diff with its callers, requirements, and evidence."},{"title":"Challenge","detail":"Exercise a case that could falsify the main claim."}],"gate":"Are material findings resolved on the current revision?","pass":"Yes → record the decision and required release checks.","fail":"No → return a reproducible finding; review the resulting change."},"sources":["google"]},{"id":"risk","title":"Route by consequence, not by author or line count","paragraphs":["A ten-line authorization change can carry more risk than a thousand-line generated fixture. Classify the consequence of being wrong, the reach of the affected code, and how easily the behavior can be observed or reversed. Record why the route was chosen. Labels such as “small” or “AI-generated” are weak substitutes for this analysis.","A low-risk copy change may need a visual check, accessibility sanity check, and content owner review. A payment calculation needs boundary examples, domain expertise, and verification of rounding and retry behavior. A migration needs compatibility and recovery analysis. The policy should explain which evidence is mandatory for each category without forcing every patch through the most expensive path.","Automation can propose the category, but ambiguous boundary crossings need a named owner. A package update that changes an authentication library is not automatically routine maintenance. If the scope expands during implementation, reroute the review rather than keeping the convenient original label."],"table":{"columns":["Change surface","Question to challenge","Useful evidence"],"rows":[["Authorization or tenant data","Can the caller cross an ownership boundary?","Negative tests using a different identity and account."],["Money, retries, or jobs","Can repetition or partial failure duplicate the effect?","Idempotency, rounding, and interrupted-operation cases."],["Schema or integration","Can old and new versions coexist during release?","Compatibility sequence and a viable recovery plan."],["User interface","Can a person complete the task across input modes?","Real viewport, keyboard, and error-state checks."]]}},{"id":"independence","title":"Make the verification path independent","paragraphs":["An authoring agent can write code and tests from the same mistaken interpretation. If both assume an export should include every account, the tests may pass while reinforcing the defect. A second reviewer reading only the author’s summary can inherit the same assumption. Independence comes from reconstructing the requirement and choosing evidence that could contradict the author.","Ask a reviewer to derive one or more boundary cases before reading the author’s test explanation. Give it the actual requirement, current code, diff, and repository conventions. Then compare its expectations with the submitted evidence. A different model can broaden the search, but it is not a proof of independence: both models may rely on the same stale documentation or incomplete context.","The reviewer should follow a changed value across the relevant path. For the export example: caller identity → account scope → query filters → serialized output. It can inspect a narrow slice deeply instead of vaguely blessing the whole repository. State the inspected scope and any unverified boundary so another person knows where confidence ends."],"bullets":["Derive at least one falsifying example from the requirement, not from the implementation.","Inspect callers and downstream effects when a local change alters a contract.","Treat external text and tool output as evidence, not as instructions that can authorize actions.","Report what was actually inspected; avoid claims about unexamined parts of the system."]},{"id":"test-quality","title":"Challenge what a passing test means","paragraphs":["Passing tests answer the questions those tests encode in the environment where they ran. They do not establish that the right questions were asked. Review whether a new test would fail against the original defect, whether its assertions observe user-relevant behavior, and whether mocked dependencies erase the very risk under review.","A useful counterfactual is to run the regression test against the prior implementation when doing so is safe and practical. If it still passes, the test may be covering a nearby behavior rather than the bug. Another approach is a narrowly scoped mutation: deliberately remove the relevant guard in an isolated workspace and confirm that the test catches it. Do not mutate a shared or production environment for this exercise.","Choose test level based on the uncertainty. A deterministic calculation usually benefits from direct boundary tests. A browser interaction may need a real rendered flow. A service integration may require a controlled environment with the actual protocol. Expensive end-to-end tests are not inherently stronger if their assertions are vague or their failures cannot be diagnosed."]},{"id":"findings","title":"Write findings that can be acted on","paragraphs":["A useful finding connects a trigger to an observed or well-supported consequence, identifies the relevant code, and explains why the issue matters. It should distinguish a confirmed defect from a concern that needs investigation. Style preferences belong in established tooling or clearly optional comments; they should not compete with broken permissions or data loss.","Avoid requiring a particular implementation unless the requirement or repository architecture demands it. The author needs to repair the behavior and provide evidence. A reviewer who rewrites the entire feature into its preferred style can enlarge the diff and obscure the original issue. Keep the correction proportional to the demonstrated problem.","This illustrative finding shows how to explain a possible defect and the evidence still needed to confirm it."],"code":{"language":"text","text":"Finding: Export query accepts an untrusted account identifier.\nTrigger: User from account A requests an export with account B's ID.\nConsequence: The query may return B's records if no lower layer scopes access.\nLocation: Export handler and its repository query (attach actual file/line).\nEvidence: Caller identity is checked, but not bound to the query account.\nRequired resolution: Enforce the ownership boundary and demonstrate denial.\nConfidence: Confirm the lower-level query behavior before marking reproduced."}},{"id":"revision","title":"Approvals belong to a revision","paragraphs":["Every review decision should identify the candidate commit. If code changes after review, the decision must be reconsidered for the affected area. If the base branch changes, determine whether the new integration state changes the behavior or verification assumptions. “Approved yesterday” is not enough information for a release system.","Define invalidation rules that are practical rather than ceremonial. A typo correction in a comment may not need a new database compatibility exercise. A lockfile change can alter runtime behavior without changing application source. The release gate should know which checks depend on code, dependencies, generated artifacts, configuration, and environment.","The reviewer’s final note should separate findings from release readiness. It can say that no material issues were found within a stated scope while still noting that a required integration environment was unavailable. Missing evidence should remain visible until an accountable person accepts the limitation or completes the check."]},{"id":"capacity","title":"Measure review as part of delivery","paragraphs":["Track how long review-ready changes wait, how many are returned for substantive corrections, and which defects escape. Do not rank reviewers by comment count or agents by approval rate. Those incentives favor noise or easy tasks. Sample accepted changes periodically to find blind spots, and use incidents to add targeted examples to the review process.","When review becomes the bottleneck, reduce the arrival rate or improve task boundaries before adding more authoring agents. Changes that each address one clear behavior can be easier to verify; arbitrarily splitting one coupled feature into many pull requests can make it harder. The useful unit is a change a reviewer can understand and accept with a bounded amount of context.","Use the downloadable code-review skill as a starting procedure. It asks for the actual diff, independent checks, revision identity, and actionable findings. It cannot certify security or replace the accountable engineer. Its value is in making a repeatable review standard explicit enough to inspect and improve."]}],"sources":[{"id":"google","title":"What to look for in a code review","author":"Google Engineering Practices","url":"https://google.github.io/eng-practices/review/reviewer/looking-for.html","kind":"Engineering guidance"}],"url":"https://vibehaus.team/blog/code-review-for-agents","markdownUrl":"https://vibehaus.team/blog/code-review-for-agents/article.md","markdown":"# Code review after agents: verify the change, not the explanation\n\nBy Vibe Haus · Published and source-reviewed 2026-10-10\n\nSource: https://vibehaus.team/blog/code-review-for-agents\n\nAn evidence-first review workflow for agent-written code, covering risk routing, independent tests, actionable findings, and exact-revision release gates.\n\nOriginal Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed.\n\n> Agent-written code deserves the same engineering standard as any other change. Faster generation makes review capacity, evidence quality, and risk selection more important.\n\n## Review three things: intent, implementation, evidence\n\nA convincing pull request description can make a change feel understood before the reviewer has read the code. Start by separating three things. Intent says what the user needs and what must remain true. Implementation is the actual diff and its surrounding behavior. Evidence shows what was checked against that intent. Agreement between all three is stronger than a polished explanation of any one.\n\nGoogle’s engineering review guidance covers design, functionality, complexity, tests, naming, comments, style, and documentation. Agent involvement does not remove those concerns. Our proposed workflow changes the order: establish the risk and the behavioral claim first, then spend review attention where a plausible patch could still be wrong.\n\nFor example, an agent may add a permission check to a new export route and describe the route as secure. The reviewer still needs to determine whose account the query uses, whether the caller can choose another account identifier, and whether the underlying data access enforces the same boundary. A correct-looking guard is not evidence that the entire path respects it.\n\n### Workflow: Review attention follows the possible consequence of the change.\n\n1. Route risk: Identify data, permissions, money, availability, and user-visible behavior.\n2. Inspect: Read the diff with its callers, requirements, and evidence.\n3. Challenge: Exercise a case that could falsify the main claim.\n\nDecision: Are material findings resolved on the current revision?\n- Yes → record the decision and required release checks.\n- No → return a reproducible finding; review the resulting change.\n\nSources: [What to look for in a code review](https://google.github.io/eng-practices/review/reviewer/looking-for.html)\n\n## Route by consequence, not by author or line count\n\nA ten-line authorization change can carry more risk than a thousand-line generated fixture. Classify the consequence of being wrong, the reach of the affected code, and how easily the behavior can be observed or reversed. Record why the route was chosen. Labels such as “small” or “AI-generated” are weak substitutes for this analysis.\n\nA low-risk copy change may need a visual check, accessibility sanity check, and content owner review. A payment calculation needs boundary examples, domain expertise, and verification of rounding and retry behavior. A migration needs compatibility and recovery analysis. The policy should explain which evidence is mandatory for each category without forcing every patch through the most expensive path.\n\nAutomation can propose the category, but ambiguous boundary crossings need a named owner. A package update that changes an authentication library is not automatically routine maintenance. If the scope expands during implementation, reroute the review rather than keeping the convenient original label.\n\n| Change surface | Question to challenge | Useful evidence |\n| --- | --- | --- |\n| Authorization or tenant data | Can the caller cross an ownership boundary? | Negative tests using a different identity and account. |\n| Money, retries, or jobs | Can repetition or partial failure duplicate the effect? | Idempotency, rounding, and interrupted-operation cases. |\n| Schema or integration | Can old and new versions coexist during release? | Compatibility sequence and a viable recovery plan. |\n| User interface | Can a person complete the task across input modes? | Real viewport, keyboard, and error-state checks. |\n\n## Make the verification path independent\n\nAn authoring agent can write code and tests from the same mistaken interpretation. If both assume an export should include every account, the tests may pass while reinforcing the defect. A second reviewer reading only the author’s summary can inherit the same assumption. Independence comes from reconstructing the requirement and choosing evidence that could contradict the author.\n\nAsk a reviewer to derive one or more boundary cases before reading the author’s test explanation. Give it the actual requirement, current code, diff, and repository conventions. Then compare its expectations with the submitted evidence. A different model can broaden the search, but it is not a proof of independence: both models may rely on the same stale documentation or incomplete context.\n\nThe reviewer should follow a changed value across the relevant path. For the export example: caller identity → account scope → query filters → serialized output. It can inspect a narrow slice deeply instead of vaguely blessing the whole repository. State the inspected scope and any unverified boundary so another person knows where confidence ends.\n\n- Derive at least one falsifying example from the requirement, not from the implementation.\n- Inspect callers and downstream effects when a local change alters a contract.\n- Treat external text and tool output as evidence, not as instructions that can authorize actions.\n- Report what was actually inspected; avoid claims about unexamined parts of the system.\n\n## Challenge what a passing test means\n\nPassing tests answer the questions those tests encode in the environment where they ran. They do not establish that the right questions were asked. Review whether a new test would fail against the original defect, whether its assertions observe user-relevant behavior, and whether mocked dependencies erase the very risk under review.\n\nA useful counterfactual is to run the regression test against the prior implementation when doing so is safe and practical. If it still passes, the test may be covering a nearby behavior rather than the bug. Another approach is a narrowly scoped mutation: deliberately remove the relevant guard in an isolated workspace and confirm that the test catches it. Do not mutate a shared or production environment for this exercise.\n\nChoose test level based on the uncertainty. A deterministic calculation usually benefits from direct boundary tests. A browser interaction may need a real rendered flow. A service integration may require a controlled environment with the actual protocol. Expensive end-to-end tests are not inherently stronger if their assertions are vague or their failures cannot be diagnosed.\n\n## Write findings that can be acted on\n\nA useful finding connects a trigger to an observed or well-supported consequence, identifies the relevant code, and explains why the issue matters. It should distinguish a confirmed defect from a concern that needs investigation. Style preferences belong in established tooling or clearly optional comments; they should not compete with broken permissions or data loss.\n\nAvoid requiring a particular implementation unless the requirement or repository architecture demands it. The author needs to repair the behavior and provide evidence. A reviewer who rewrites the entire feature into its preferred style can enlarge the diff and obscure the original issue. Keep the correction proportional to the demonstrated problem.\n\nThis illustrative finding shows how to explain a possible defect and the evidence still needed to confirm it.\n\n```text\nFinding: Export query accepts an untrusted account identifier.\nTrigger: User from account A requests an export with account B's ID.\nConsequence: The query may return B's records if no lower layer scopes access.\nLocation: Export handler and its repository query (attach actual file/line).\nEvidence: Caller identity is checked, but not bound to the query account.\nRequired resolution: Enforce the ownership boundary and demonstrate denial.\nConfidence: Confirm the lower-level query behavior before marking reproduced.\n```\n\n## Approvals belong to a revision\n\nEvery review decision should identify the candidate commit. If code changes after review, the decision must be reconsidered for the affected area. If the base branch changes, determine whether the new integration state changes the behavior or verification assumptions. “Approved yesterday” is not enough information for a release system.\n\nDefine invalidation rules that are practical rather than ceremonial. A typo correction in a comment may not need a new database compatibility exercise. A lockfile change can alter runtime behavior without changing application source. The release gate should know which checks depend on code, dependencies, generated artifacts, configuration, and environment.\n\nThe reviewer’s final note should separate findings from release readiness. It can say that no material issues were found within a stated scope while still noting that a required integration environment was unavailable. Missing evidence should remain visible until an accountable person accepts the limitation or completes the check.\n\n## Measure review as part of delivery\n\nTrack how long review-ready changes wait, how many are returned for substantive corrections, and which defects escape. Do not rank reviewers by comment count or agents by approval rate. Those incentives favor noise or easy tasks. Sample accepted changes periodically to find blind spots, and use incidents to add targeted examples to the review process.\n\nWhen review becomes the bottleneck, reduce the arrival rate or improve task boundaries before adding more authoring agents. Changes that each address one clear behavior can be easier to verify; arbitrarily splitting one coupled feature into many pull requests can make it harder. The useful unit is a change a reviewer can understand and accept with a bounded amount of context.\n\nUse the downloadable code-review skill as a starting procedure. It asks for the actual diff, independent checks, revision identity, and actionable findings. It cannot certify security or replace the accountable engineer. Its value is in making a repeatable review standard explicit enough to inspect and improve.\n\n## Sources and further reading\n\nSources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations.\n\n- [What to look for in a code review](https://google.github.io/eng-practices/review/reviewer/looking-for.html): Google Engineering Practices. Engineering guidance.\n\n## Related services\n\n- [ai native engineering](https://vibehaus.team/ai-native-engineering)\n\n- [managed engineering](https://vibehaus.team/managed-engineering)\n\n## Reuse\n\nYou may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms.\n\nDownloads and tools: https://vibehaus.team/data"},{"slug":"context-mcp-and-skills","title":"Context engineering: repository knowledge, MCP, and skills","description":"How repository knowledge, task context, MCP tools, and agent skills fit together, with provenance, freshness, trust boundaries, and a practical context contract.","topic":"Context & tools","date":"2026-10-10","thesis":"Agents need relevant, current information for the task they are doing. This guide explains how to select that information, provide tools and procedures, and maintain the underlying sources.","sections":[{"id":"layers","title":"Separate knowledge, procedures, and access","paragraphs":["An engineering agent needs several different things that are easy to collapse into one large prompt. It needs durable knowledge about the product and repository, temporary facts about the current task, procedures for recurring work, and access to tools or external systems. These layers change at different rates and carry different authority.","A repository guide can describe where code lives and which checks are expected. A task contract can identify today’s desired behavior. A skill can explain how to perform a recurring review. The Model Context Protocol (MCP) gives agents a standard way to call tools, such as search or document retrieval. None of those, by itself, establishes that an external document is accurate or that a user authorized a consequential action.","The Agent Skills specification defines a discoverable SKILL.md with a name and description, with additional material loaded when useful. MCP defines a protocol for exposing capabilities and exchanging messages. These are complementary mechanisms: a procedure can use a tool, while the tool should retain its own access controls and input validation."],"flow":{"caption":"A context assembly path with a freshness and authority check.","steps":[{"title":"Locate","detail":"Start with the repository map and task contract."},{"title":"Retrieve","detail":"Read the smallest relevant source or tool result."},{"title":"Qualify","detail":"Record provenance, revision, and unresolved conflicts."}],"gate":"Is the context sufficient and trustworthy for this decision?","pass":"Yes → act within the existing authority and preserve evidence.","fail":"No → retrieve the missing fact or return the unresolved decision."},"sources":["skills","mcp"]},{"id":"map","title":"Keep repository instructions concise and current","paragraphs":["A repository map should answer where to look next. Name the entry points, ownership boundaries, local commands, and important architectural decisions. Link to detailed documents instead of copying them into a single instruction file. The goal is to help an agent locate the relevant source of truth and recognize when the task crosses into another area.","Dex Horthy’s AI Engineer talk presents a research, planning, and implementation approach for complex codebases. We take a narrow lesson from that framing: understanding the current system is a separate artifact worth checking before a broad edit. A plausible plan built on the wrong version of the code remains a bad plan.","For a billing task, the map might lead to the account model, invoice calculation, external payment adapter, and migration policy. It should also distinguish maintained decisions from old exploration notes. A workshop proposal and an accepted architecture decision can both be useful, but treating them as equally authoritative produces avoidable contradictions."],"bullets":["Entry points: the few paths that orient a reader to the feature.","Ownership: who resolves product, data, and release decisions.","Verification: commands and environments that check the relevant behavior.","Status: accepted decisions, current implementation, and explicitly tentative ideas."],"sources":["horthy"]},{"id":"task-context","title":"Build a context contract for the current task","paragraphs":["Task context should explain the decision being made, the version of the system under discussion, and the boundaries that matter. A file path without a revision can become misleading after a rename or refactor. A log excerpt without environment and timestamp can describe a different deployment. Preserve enough provenance to reconnect a claim to the thing it describes.","The contract can remain compact. Include the outcome, relevant source paths, current revision, constraints, acceptance examples, known unknowns, and the next decision. Separate observed facts from hypotheses. “The export job times out in staging with this fixture” is an observation. “The database needs an index” is a hypothesis until examined.","When the work moves to another agent or person, summarize what was learned and where the supporting evidence lives. Do not simply compress every prior message. Preserve decisions, rejected alternatives that still matter, current failures, and unresolved questions. A handoff should let the recipient resume reasoning without inheriting an unexamined conclusion."],"code":{"language":"yaml","text":"# Illustrative context contract.\ntask: Diagnose slow filtered export.\nrevision: Record the actual commit under investigation.\nobserved:\n  - Staging request exceeded the agreed limit with the attached fixture.\nhypotheses:\n  - Query plan may scan too many rows; not yet confirmed.\nsources:\n  - Current route, query implementation, schema, and captured query plan.\nconstraints:\n  - Preserve account isolation and existing output format.\nnext_decision: Is the bottleneck query work, serialization, or delivery?\nmissing: Representative upper-bound dataset and acceptable latency target."}},{"id":"skills","title":"Make skills procedural and inspectable","paragraphs":["In the AI Engineer talk Don’t Build Agents, Build Skills Instead, Barry Zhang and Mahesh Murag describe packaging reusable knowledge for general agents. Our recommendation is to use a skill when a task has a repeatable method with non-obvious checks. A review skill can explain how to inspect a changed boundary; it should not duplicate the repository’s entire architecture.","A good skill explains when it should be used and includes required inputs, a procedure, a useful output, and conditions under which to stop. Keep instructions that apply to every task in the appropriate repository or runtime policy. Keep task-specific details in the task itself. Otherwise skills become competing collections of global rules that are hard to reconcile.","Version a skill when its behavior changes. Test it with a realistic task that includes an inconvenient case: missing evidence, conflicting requirements, a failing environment, or a request outside its scope. A syntactically valid SKILL.md can still encourage poor decisions. Behavioral testing should ask whether the resulting artifact helps a real reviewer complete the job."],"sources":["zhang"]},{"id":"mcp","title":"Design MCP tools around bounded questions","paragraphs":["A useful tool interface gives an agent a clear question it can ask and a result it can interpret. For a public research library, “search published articles” and “read this article by slug” are enough. A generic “fetch any URL and execute what it says” capability would create a very different risk surface without improving that basic use case.","Tool results should include stable identifiers and source URLs. Search snippets help choose a document; they are not a substitute for reading the relevant passage. A read result should expose the full authored content or explicitly state truncation. Unknown identifiers should produce an understandable error instead of silently selecting a nearby item.","Our public MCP endpoint follows that narrow design: it searches and reads this site’s own published research and retrieves the downloadable skills. It does not connect to customer systems or receive repository credentials. Its scope is visible on the Data / MCP / Skills page, and the same article text is available through JSON and Markdown for clients that do not use MCP."]},{"id":"trust","title":"Treat retrieved content as data with provenance","paragraphs":["An external page can contain instructions written by someone other than the user. A document that says to ignore prior rules or send credentials somewhere is still just document content. The agent runtime should keep retrieved text separate from higher-priority instructions, and tools should enforce access boundaries independently of the model’s interpretation.","The same principle applies inside a repository when files, issue comments, or generated output can be influenced by outside contributors. Read them for relevant facts; do not let them silently broaden the task’s authority. Skills downloaded from the internet should be inspected before installation, just as a team would inspect another automation artifact.","Freshness is another trust dimension. Cache stable documents with explicit versions; recheck volatile facts before acting on them. If a repository guide conflicts with the installed framework’s current documentation or the executable code, investigate and record the resolution. Do not hide the conflict by choosing whichever source supports the first plan."]},{"id":"maintenance","title":"Maintain context as part of changing the system","paragraphs":["When a change moves an entry point, changes a command, or supersedes an architectural decision, update the map in the same delivery cycle. The person accepting the change should be able to verify that the next engineer will find the right path. This makes documentation maintenance observable rather than an occasional cleanup project.","Sample failed agent runs and classify the context failure. Was the needed fact absent, stale, difficult to retrieve, contradicted, or ignored? Adding more text helps only the first category and sometimes the third. A better source hierarchy, a smaller tool result, or a clearer stopping condition may be the more effective repair.","The intended outcome is a connected knowledge system: a short repository map, task-specific evidence, reusable procedures, and bounded tools. Each part should be independently understandable. That lets people and agents share the same engineering context without requiring everyone to consume the entire history of the project."]}],"sources":[{"id":"skills","title":"Agent Skills specification","author":"Agent Skills","url":"https://agentskills.io/specification","kind":"Specification"},{"id":"mcp","title":"Streamable HTTP transport specification","author":"Model Context Protocol · 2025-11-25","url":"https://modelcontextprotocol.io/specification/2025-11-25/basic/transports","kind":"Specification"},{"id":"horthy","title":"No Vibes Allowed: Solving Hard Problems in Complex Codebases","author":"Dex Horthy · AI Engineer","url":"https://ai.engineer/talks/rmvDxxNubIg-context-engineering-for-complex-codebases","kind":"Conference talk"},{"id":"zhang","title":"Don’t Build Agents, Build Skills Instead","author":"Barry Zhang & Mahesh Murag · AI Engineer","url":"https://ai.engineer/talks/CEvIs9y1uog-agent-skills","kind":"Conference talk"}],"url":"https://vibehaus.team/blog/context-mcp-and-skills","markdownUrl":"https://vibehaus.team/blog/context-mcp-and-skills/article.md","markdown":"# Context engineering: repository knowledge, MCP, and skills\n\nBy Vibe Haus · Published and source-reviewed 2026-10-10\n\nSource: https://vibehaus.team/blog/context-mcp-and-skills\n\nHow repository knowledge, task context, MCP tools, and agent skills fit together, with provenance, freshness, trust boundaries, and a practical context contract.\n\nOriginal Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed.\n\n> Agents need relevant, current information for the task they are doing. This guide explains how to select that information, provide tools and procedures, and maintain the underlying sources.\n\n## Separate knowledge, procedures, and access\n\nAn engineering agent needs several different things that are easy to collapse into one large prompt. It needs durable knowledge about the product and repository, temporary facts about the current task, procedures for recurring work, and access to tools or external systems. These layers change at different rates and carry different authority.\n\nA repository guide can describe where code lives and which checks are expected. A task contract can identify today’s desired behavior. A skill can explain how to perform a recurring review. The Model Context Protocol (MCP) gives agents a standard way to call tools, such as search or document retrieval. None of those, by itself, establishes that an external document is accurate or that a user authorized a consequential action.\n\nThe Agent Skills specification defines a discoverable SKILL.md with a name and description, with additional material loaded when useful. MCP defines a protocol for exposing capabilities and exchanging messages. These are complementary mechanisms: a procedure can use a tool, while the tool should retain its own access controls and input validation.\n\n### Workflow: A context assembly path with a freshness and authority check.\n\n1. Locate: Start with the repository map and task contract.\n2. Retrieve: Read the smallest relevant source or tool result.\n3. Qualify: Record provenance, revision, and unresolved conflicts.\n\nDecision: Is the context sufficient and trustworthy for this decision?\n- Yes → act within the existing authority and preserve evidence.\n- No → retrieve the missing fact or return the unresolved decision.\n\nSources: [Agent Skills specification](https://agentskills.io/specification); [Streamable HTTP transport specification](https://modelcontextprotocol.io/specification/2025-11-25/basic/transports)\n\n## Keep repository instructions concise and current\n\nA repository map should answer where to look next. Name the entry points, ownership boundaries, local commands, and important architectural decisions. Link to detailed documents instead of copying them into a single instruction file. The goal is to help an agent locate the relevant source of truth and recognize when the task crosses into another area.\n\nDex Horthy’s AI Engineer talk presents a research, planning, and implementation approach for complex codebases. We take a narrow lesson from that framing: understanding the current system is a separate artifact worth checking before a broad edit. A plausible plan built on the wrong version of the code remains a bad plan.\n\nFor a billing task, the map might lead to the account model, invoice calculation, external payment adapter, and migration policy. It should also distinguish maintained decisions from old exploration notes. A workshop proposal and an accepted architecture decision can both be useful, but treating them as equally authoritative produces avoidable contradictions.\n\n- Entry points: the few paths that orient a reader to the feature.\n- Ownership: who resolves product, data, and release decisions.\n- Verification: commands and environments that check the relevant behavior.\n- Status: accepted decisions, current implementation, and explicitly tentative ideas.\n\nSources: [No Vibes Allowed: Solving Hard Problems in Complex Codebases](https://ai.engineer/talks/rmvDxxNubIg-context-engineering-for-complex-codebases)\n\n## Build a context contract for the current task\n\nTask context should explain the decision being made, the version of the system under discussion, and the boundaries that matter. A file path without a revision can become misleading after a rename or refactor. A log excerpt without environment and timestamp can describe a different deployment. Preserve enough provenance to reconnect a claim to the thing it describes.\n\nThe contract can remain compact. Include the outcome, relevant source paths, current revision, constraints, acceptance examples, known unknowns, and the next decision. Separate observed facts from hypotheses. “The export job times out in staging with this fixture” is an observation. “The database needs an index” is a hypothesis until examined.\n\nWhen the work moves to another agent or person, summarize what was learned and where the supporting evidence lives. Do not simply compress every prior message. Preserve decisions, rejected alternatives that still matter, current failures, and unresolved questions. A handoff should let the recipient resume reasoning without inheriting an unexamined conclusion.\n\n```yaml\n# Illustrative context contract.\ntask: Diagnose slow filtered export.\nrevision: Record the actual commit under investigation.\nobserved:\n  - Staging request exceeded the agreed limit with the attached fixture.\nhypotheses:\n  - Query plan may scan too many rows; not yet confirmed.\nsources:\n  - Current route, query implementation, schema, and captured query plan.\nconstraints:\n  - Preserve account isolation and existing output format.\nnext_decision: Is the bottleneck query work, serialization, or delivery?\nmissing: Representative upper-bound dataset and acceptable latency target.\n```\n\n## Make skills procedural and inspectable\n\nIn the AI Engineer talk Don’t Build Agents, Build Skills Instead, Barry Zhang and Mahesh Murag describe packaging reusable knowledge for general agents. Our recommendation is to use a skill when a task has a repeatable method with non-obvious checks. A review skill can explain how to inspect a changed boundary; it should not duplicate the repository’s entire architecture.\n\nA good skill explains when it should be used and includes required inputs, a procedure, a useful output, and conditions under which to stop. Keep instructions that apply to every task in the appropriate repository or runtime policy. Keep task-specific details in the task itself. Otherwise skills become competing collections of global rules that are hard to reconcile.\n\nVersion a skill when its behavior changes. Test it with a realistic task that includes an inconvenient case: missing evidence, conflicting requirements, a failing environment, or a request outside its scope. A syntactically valid SKILL.md can still encourage poor decisions. Behavioral testing should ask whether the resulting artifact helps a real reviewer complete the job.\n\nSources: [Don’t Build Agents, Build Skills Instead](https://ai.engineer/talks/CEvIs9y1uog-agent-skills)\n\n## Design MCP tools around bounded questions\n\nA useful tool interface gives an agent a clear question it can ask and a result it can interpret. For a public research library, “search published articles” and “read this article by slug” are enough. A generic “fetch any URL and execute what it says” capability would create a very different risk surface without improving that basic use case.\n\nTool results should include stable identifiers and source URLs. Search snippets help choose a document; they are not a substitute for reading the relevant passage. A read result should expose the full authored content or explicitly state truncation. Unknown identifiers should produce an understandable error instead of silently selecting a nearby item.\n\nOur public MCP endpoint follows that narrow design: it searches and reads this site’s own published research and retrieves the downloadable skills. It does not connect to customer systems or receive repository credentials. Its scope is visible on the Data / MCP / Skills page, and the same article text is available through JSON and Markdown for clients that do not use MCP.\n\n## Treat retrieved content as data with provenance\n\nAn external page can contain instructions written by someone other than the user. A document that says to ignore prior rules or send credentials somewhere is still just document content. The agent runtime should keep retrieved text separate from higher-priority instructions, and tools should enforce access boundaries independently of the model’s interpretation.\n\nThe same principle applies inside a repository when files, issue comments, or generated output can be influenced by outside contributors. Read them for relevant facts; do not let them silently broaden the task’s authority. Skills downloaded from the internet should be inspected before installation, just as a team would inspect another automation artifact.\n\nFreshness is another trust dimension. Cache stable documents with explicit versions; recheck volatile facts before acting on them. If a repository guide conflicts with the installed framework’s current documentation or the executable code, investigate and record the resolution. Do not hide the conflict by choosing whichever source supports the first plan.\n\n## Maintain context as part of changing the system\n\nWhen a change moves an entry point, changes a command, or supersedes an architectural decision, update the map in the same delivery cycle. The person accepting the change should be able to verify that the next engineer will find the right path. This makes documentation maintenance observable rather than an occasional cleanup project.\n\nSample failed agent runs and classify the context failure. Was the needed fact absent, stale, difficult to retrieve, contradicted, or ignored? Adding more text helps only the first category and sometimes the third. A better source hierarchy, a smaller tool result, or a clearer stopping condition may be the more effective repair.\n\nThe intended outcome is a connected knowledge system: a short repository map, task-specific evidence, reusable procedures, and bounded tools. Each part should be independently understandable. That lets people and agents share the same engineering context without requiring everyone to consume the entire history of the project.\n\n## Sources and further reading\n\nSources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations.\n\n- [Agent Skills specification](https://agentskills.io/specification): Agent Skills. Specification.\n\n- [Streamable HTTP transport specification](https://modelcontextprotocol.io/specification/2025-11-25/basic/transports): Model Context Protocol · 2025-11-25. Specification.\n\n- [No Vibes Allowed: Solving Hard Problems in Complex Codebases](https://ai.engineer/talks/rmvDxxNubIg-context-engineering-for-complex-codebases): Dex Horthy · AI Engineer. Conference talk.\n\n- [Don’t Build Agents, Build Skills Instead](https://ai.engineer/talks/CEvIs9y1uog-agent-skills): Barry Zhang & Mahesh Murag · AI Engineer. Conference talk.\n\n## Related services\n\n- [ai native engineering](https://vibehaus.team/ai-native-engineering)\n\n- [product design](https://vibehaus.team/product-design)\n\n## Reuse\n\nYou may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms.\n\nDownloads and tools: https://vibehaus.team/data"},{"slug":"evaluating-agentic-engineering","title":"Evaluating agentic engineering: quality, cost, and delivery","description":"A measurement plan for agentic engineering that includes accepted outcomes, human review, rework, failures, cost, and the limits of current productivity evidence.","topic":"Evaluation","date":"2026-10-10","thesis":"The question is whether the team delivers more valuable, dependable work for the resources it spends. Token speed, generated lines, and convincing demos cannot answer that alone.","sections":[{"id":"evidence","title":"Read productivity claims in their actual setting","paragraphs":["METR’s early-2025 randomized study followed 16 experienced open-source developers working on 246 tasks in repositories they knew well. With the tested AI tools available, tasks took 19% longer on average. This is useful evidence about that population, task mix, and tool period. It is not a universal estimate for every developer, greenfield product, or later agent.","The February 2026 follow-up matters too. METR reports that selection effects and difficulties measuring concurrent agent work make its newer productivity estimates unreliable. Developers and tasks opting out of the no-AI condition can change who remains in the experiment. The authors suggest improvements are plausible while warning that their data gives weak evidence for the size of the change.","DORA’s 2025 report frames AI as amplifying existing organizational strengths and weaknesses. Treat that organizational research as context for adoption, not as a causal guarantee for your team. The practical response to mixed and changing evidence is to measure your own work carefully, while keeping the limits of that measurement visible."],"sources":["metr2025","metr2026","dora"]},{"id":"unit","title":"Define the work being measured","paragraphs":["A completed coding task is not necessarily a delivered product outcome. Decide whether the experiment concerns implementation speed, accepted changes, released features, or customer results. Each requires a different observation window. A benchmark patch can be evaluated before release; customer retention cannot.","For an initial engineering pilot, use a change with a clear outcome that a reviewer can accept. Define completion in advance: required behavior, checks, review disposition, and unresolved limitations. Include failed and abandoned attempts in the denominator. Otherwise a system that spends heavily on ten attempts and succeeds once can appear equivalent to one that succeeds immediately.","Segment the task set. A documentation correction, an unfamiliar integration, and a concurrency bug exercise different capabilities. Record repository familiarity, task uncertainty, risk, and dependency complexity. Report performance by meaningful segment before combining it into an average. A gain on easy tasks can conceal a regression on the work that occupies most of the team."],"flow":{"caption":"A proposed evaluation loop. Keep a held-out set separate from workflow tuning.","steps":[{"title":"Define","detail":"Choose tasks, acceptance rules, and the decision to be made."},{"title":"Compare","detail":"Run a baseline and candidate under recorded conditions."},{"title":"Inspect","detail":"Score quality, total effort, failures, and cost."}],"gate":"Does the candidate meet the predeclared quality and cost criteria?","pass":"Yes → expand gradually and keep monitoring production outcomes.","fail":"No → diagnose by task segment, change one cause, then reevaluate."}},{"id":"dataset","title":"Build a small, representative task set","paragraphs":["Start with real task shapes from the team’s backlog or recent history. Historical tasks are useful only if the evaluation prevents the candidate from seeing the finished solution. Freeze the starting repository revision, remove answer-bearing artifacts from the candidate’s context, and check whether the model or tool could retrieve the solution elsewhere. Report contamination you cannot exclude.","Write acceptance criteria independently of the candidate implementation. Include ordinary paths, relevant boundary cases, and at least some tasks that require recognizing insufficient information. An agent that stops for a missing authorization decision may be behaving better than one that produces a complete-looking patch by inventing the answer.","Keep a development set for tuning prompts, tools, and skills, and a separate held-out set for decisions. Repeatedly optimizing against the same examples can produce a workflow that is good at the evaluation rather than the job. Version the tasks and record when a task becomes unsuitable because the product or tooling changed."],"bullets":["Preserve a reproducible starting state and the available context.","Specify acceptance evidence before seeing the candidate result.","Include failures, ambiguity, and realistic constraints.","Keep held-out tasks out of routine prompt and skill tuning."]},{"id":"comparison","title":"Compare workflows and document the limits","paragraphs":["A baseline can be the team’s current workflow, a simpler agent setup, or a different model configuration. Choose the comparison that answers the adoption decision. If the question is whether an orchestration layer helps, hold the model and task conditions as stable as practical instead of changing everything at once.","For task-level comparisons, randomize assignment where feasible and balance important task categories. Repeating the same task with the same developer introduces learning; paired tasks need to be comparable rather than identical in a way that leaks the solution. Track deviations from the assigned workflow and report exclusions explicitly. Small pilots provide directional evidence, not a precise universal effect.","Agents introduce variability. Record model identifier, tool versions, reasoning settings when available, skills, repository revision, resource limits, and number of attempts. A single successful run cannot establish reliability. Repeat enough cases to see whether the result depends on a lucky attempt, and report the distribution rather than only the best outcome."]},{"id":"cost","title":"Count the work that generation metrics omit","paragraphs":["Capture human time spent framing, supervising, reviewing, integrating, and repairing. Capture model, tool, and infrastructure cost separately. Record wall-clock lead time as well: a task can consume little active human time while waiting several days in a review queue. Parallel agents make these measures diverge even more.","Use accepted outcomes as the denominator for an economic comparison, while keeping quality thresholds fixed. The illustrative example below assumes the same task mix and acceptance standard. It is arithmetic to show the method, not a measured Vibe Haus result or a promise about an agent system.","Here the candidate uses less total money but produces fewer accepted changes. Cost per accepted change therefore rises from $100 to about $105.56. The result could still be acceptable for another reason, such as shorter lead time, but that reason must be measured rather than inferred from the smaller total bill."],"code":{"language":"text","text":"cost per accepted change =\n  (human effort cost + model cost + tools + infrastructure)\n  / accepted changes\n\nReport separately:\n- unresolved rework and the observation window\n- quality failures and severity\n- request-to-acceptance lead time\n- tasks excluded or abandoned"},"table":{"columns":["Illustrative measure","Baseline","Candidate"],"rows":[["Attempted changes","12","12"],["Accepted changes","10","9"],["Human effort at $100/hour","9 hours = $900","7 hours = $700"],["Model, tools, infrastructure","$100","$250"],["Total measured cost","$1,000","$950"],["Cost per accepted change","$100","$105.56"]]}},{"id":"quality","title":"Define acceptance thresholds before the trial","paragraphs":["Predeclare the conditions that would stop or limit rollout. These might include an unacceptable permission defect, a cost ceiling, excessive reviewer intervention, or unreliable completion on a critical task family. The thresholds belong to the product’s risk and economics. Do not choose them after seeing the results to make a preferred tool win.","Combine deterministic checks with independent judgment where necessary. Tests can establish explicit behavior; expert review can examine architecture, maintainability, and missing cases. Reviewers should use a shared rubric and record disagreements. When practical, hide which workflow produced a candidate to reduce expectation effects, while recognizing that style or artifacts may reveal it.","Track escaped defects over an appropriate window after acceptance. A pilot that ends at merge cannot measure downstream maintenance or incidents. Keep those outcomes separate from immediate pass rate, and describe the lag. The absence of an observed incident in a small short trial is weak evidence about rare high-impact failures."]},{"id":"decision","title":"Turn the result into an operating decision","paragraphs":["An evaluation should end with a specific decision: adopt for a named task family, continue a bounded pilot, revise the workflow, or stop. Explain what evidence would change that decision. A partial success can justify a narrow deployment while leaving higher-risk work under a different process.","If the candidate fails, diagnose the stage. Ambiguous tasks call for better contracts. Repeated tool misuse calls for better interfaces or constraints. High review effort may indicate broad diffs or weak evidence. Expensive retries may point to an unsuitable task class. Changing the model is one possible intervention, not the default explanation for every failure.","The downloadable evaluation skill turns this article into a reusable plan and report structure. Use it before a pilot to make the comparison fair, then again afterward to identify missing data and limitations. The aim is a decision the team can defend with observed outcomes, including the inconvenient ones."]}],"sources":[{"id":"metr2025","title":"Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity","author":"METR · July 10, 2025","url":"https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/","kind":"Randomized study"},{"id":"metr2026","title":"We are Changing our Developer Productivity Experiment Design","author":"METR · February 24, 2026","url":"https://metr.org/blog/2026-02-24-uplift-update/","kind":"Study update"},{"id":"dora","title":"State of AI-assisted Software Development","author":"DORA · 2025","url":"https://dora.dev/research/2025/dora-report/","kind":"Research report"}],"url":"https://vibehaus.team/blog/evaluating-agentic-engineering","markdownUrl":"https://vibehaus.team/blog/evaluating-agentic-engineering/article.md","markdown":"# Evaluating agentic engineering: quality, cost, and delivery\n\nBy Vibe Haus · Published and source-reviewed 2026-10-10\n\nSource: https://vibehaus.team/blog/evaluating-agentic-engineering\n\nA measurement plan for agentic engineering that includes accepted outcomes, human review, rework, failures, cost, and the limits of current productivity evidence.\n\nOriginal Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed.\n\n> The question is whether the team delivers more valuable, dependable work for the resources it spends. Token speed, generated lines, and convincing demos cannot answer that alone.\n\n## Read productivity claims in their actual setting\n\nMETR’s early-2025 randomized study followed 16 experienced open-source developers working on 246 tasks in repositories they knew well. With the tested AI tools available, tasks took 19% longer on average. This is useful evidence about that population, task mix, and tool period. It is not a universal estimate for every developer, greenfield product, or later agent.\n\nThe February 2026 follow-up matters too. METR reports that selection effects and difficulties measuring concurrent agent work make its newer productivity estimates unreliable. Developers and tasks opting out of the no-AI condition can change who remains in the experiment. The authors suggest improvements are plausible while warning that their data gives weak evidence for the size of the change.\n\nDORA’s 2025 report frames AI as amplifying existing organizational strengths and weaknesses. Treat that organizational research as context for adoption, not as a causal guarantee for your team. The practical response to mixed and changing evidence is to measure your own work carefully, while keeping the limits of that measurement visible.\n\nSources: [Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/); [We are Changing our Developer Productivity Experiment Design](https://metr.org/blog/2026-02-24-uplift-update/); [State of AI-assisted Software Development](https://dora.dev/research/2025/dora-report/)\n\n## Define the work being measured\n\nA completed coding task is not necessarily a delivered product outcome. Decide whether the experiment concerns implementation speed, accepted changes, released features, or customer results. Each requires a different observation window. A benchmark patch can be evaluated before release; customer retention cannot.\n\nFor an initial engineering pilot, use a change with a clear outcome that a reviewer can accept. Define completion in advance: required behavior, checks, review disposition, and unresolved limitations. Include failed and abandoned attempts in the denominator. Otherwise a system that spends heavily on ten attempts and succeeds once can appear equivalent to one that succeeds immediately.\n\nSegment the task set. A documentation correction, an unfamiliar integration, and a concurrency bug exercise different capabilities. Record repository familiarity, task uncertainty, risk, and dependency complexity. Report performance by meaningful segment before combining it into an average. A gain on easy tasks can conceal a regression on the work that occupies most of the team.\n\n### Workflow: A proposed evaluation loop. Keep a held-out set separate from workflow tuning.\n\n1. Define: Choose tasks, acceptance rules, and the decision to be made.\n2. Compare: Run a baseline and candidate under recorded conditions.\n3. Inspect: Score quality, total effort, failures, and cost.\n\nDecision: Does the candidate meet the predeclared quality and cost criteria?\n- Yes → expand gradually and keep monitoring production outcomes.\n- No → diagnose by task segment, change one cause, then reevaluate.\n\n## Build a small, representative task set\n\nStart with real task shapes from the team’s backlog or recent history. Historical tasks are useful only if the evaluation prevents the candidate from seeing the finished solution. Freeze the starting repository revision, remove answer-bearing artifacts from the candidate’s context, and check whether the model or tool could retrieve the solution elsewhere. Report contamination you cannot exclude.\n\nWrite acceptance criteria independently of the candidate implementation. Include ordinary paths, relevant boundary cases, and at least some tasks that require recognizing insufficient information. An agent that stops for a missing authorization decision may be behaving better than one that produces a complete-looking patch by inventing the answer.\n\nKeep a development set for tuning prompts, tools, and skills, and a separate held-out set for decisions. Repeatedly optimizing against the same examples can produce a workflow that is good at the evaluation rather than the job. Version the tasks and record when a task becomes unsuitable because the product or tooling changed.\n\n- Preserve a reproducible starting state and the available context.\n- Specify acceptance evidence before seeing the candidate result.\n- Include failures, ambiguity, and realistic constraints.\n- Keep held-out tasks out of routine prompt and skill tuning.\n\n## Compare workflows and document the limits\n\nA baseline can be the team’s current workflow, a simpler agent setup, or a different model configuration. Choose the comparison that answers the adoption decision. If the question is whether an orchestration layer helps, hold the model and task conditions as stable as practical instead of changing everything at once.\n\nFor task-level comparisons, randomize assignment where feasible and balance important task categories. Repeating the same task with the same developer introduces learning; paired tasks need to be comparable rather than identical in a way that leaks the solution. Track deviations from the assigned workflow and report exclusions explicitly. Small pilots provide directional evidence, not a precise universal effect.\n\nAgents introduce variability. Record model identifier, tool versions, reasoning settings when available, skills, repository revision, resource limits, and number of attempts. A single successful run cannot establish reliability. Repeat enough cases to see whether the result depends on a lucky attempt, and report the distribution rather than only the best outcome.\n\n## Count the work that generation metrics omit\n\nCapture human time spent framing, supervising, reviewing, integrating, and repairing. Capture model, tool, and infrastructure cost separately. Record wall-clock lead time as well: a task can consume little active human time while waiting several days in a review queue. Parallel agents make these measures diverge even more.\n\nUse accepted outcomes as the denominator for an economic comparison, while keeping quality thresholds fixed. The illustrative example below assumes the same task mix and acceptance standard. It is arithmetic to show the method, not a measured Vibe Haus result or a promise about an agent system.\n\nHere the candidate uses less total money but produces fewer accepted changes. Cost per accepted change therefore rises from $100 to about $105.56. The result could still be acceptable for another reason, such as shorter lead time, but that reason must be measured rather than inferred from the smaller total bill.\n\n```text\ncost per accepted change =\n  (human effort cost + model cost + tools + infrastructure)\n  / accepted changes\n\nReport separately:\n- unresolved rework and the observation window\n- quality failures and severity\n- request-to-acceptance lead time\n- tasks excluded or abandoned\n```\n\n| Illustrative measure | Baseline | Candidate |\n| --- | --- | --- |\n| Attempted changes | 12 | 12 |\n| Accepted changes | 10 | 9 |\n| Human effort at $100/hour | 9 hours = $900 | 7 hours = $700 |\n| Model, tools, infrastructure | $100 | $250 |\n| Total measured cost | $1,000 | $950 |\n| Cost per accepted change | $100 | $105.56 |\n\n## Define acceptance thresholds before the trial\n\nPredeclare the conditions that would stop or limit rollout. These might include an unacceptable permission defect, a cost ceiling, excessive reviewer intervention, or unreliable completion on a critical task family. The thresholds belong to the product’s risk and economics. Do not choose them after seeing the results to make a preferred tool win.\n\nCombine deterministic checks with independent judgment where necessary. Tests can establish explicit behavior; expert review can examine architecture, maintainability, and missing cases. Reviewers should use a shared rubric and record disagreements. When practical, hide which workflow produced a candidate to reduce expectation effects, while recognizing that style or artifacts may reveal it.\n\nTrack escaped defects over an appropriate window after acceptance. A pilot that ends at merge cannot measure downstream maintenance or incidents. Keep those outcomes separate from immediate pass rate, and describe the lag. The absence of an observed incident in a small short trial is weak evidence about rare high-impact failures.\n\n## Turn the result into an operating decision\n\nAn evaluation should end with a specific decision: adopt for a named task family, continue a bounded pilot, revise the workflow, or stop. Explain what evidence would change that decision. A partial success can justify a narrow deployment while leaving higher-risk work under a different process.\n\nIf the candidate fails, diagnose the stage. Ambiguous tasks call for better contracts. Repeated tool misuse calls for better interfaces or constraints. High review effort may indicate broad diffs or weak evidence. Expensive retries may point to an unsuitable task class. Changing the model is one possible intervention, not the default explanation for every failure.\n\nThe downloadable evaluation skill turns this article into a reusable plan and report structure. Use it before a pilot to make the comparison fair, then again afterward to identify missing data and limitations. The aim is a decision the team can defend with observed outcomes, including the inconvenient ones.\n\n## Sources and further reading\n\nSources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations.\n\n- [Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/): METR · July 10, 2025. Randomized study.\n\n- [We are Changing our Developer Productivity Experiment Design](https://metr.org/blog/2026-02-24-uplift-update/): METR · February 24, 2026. Study update.\n\n- [State of AI-assisted Software Development](https://dora.dev/research/2025/dora-report/): DORA · 2025. Research report.\n\n## Related services\n\n- [research and development](https://vibehaus.team/research-and-development)\n\n- [engineering cost](https://vibehaus.team/engineering-cost)\n\n## Reuse\n\nYou may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms.\n\nDownloads and tools: https://vibehaus.team/data"}]}