Read productivity claims in their actual setting

METR’s early-2025 randomized study followed 16 experienced open-source developers working on 246 tasks in repositories they knew well. With the tested AI tools available, tasks took 19% longer on average. This is useful evidence about that population, task mix, and tool period. It is not a universal estimate for every developer, greenfield product, or later agent.

The February 2026 follow-up matters too. METR reports that selection effects and difficulties measuring concurrent agent work make its newer productivity estimates unreliable. Developers and tasks opting out of the no-AI condition can change who remains in the experiment. The authors suggest improvements are plausible while warning that their data gives weak evidence for the size of the change.

DORA’s 2025 report frames AI as amplifying existing organizational strengths and weaknesses. Treat that organizational research as context for adoption, not as a causal guarantee for your team. The practical response to mixed and changing evidence is to measure your own work carefully, while keeping the limits of that measurement visible.

Sources: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · We are Changing our Developer Productivity Experiment Design · State of AI-assisted Software Development

Define the work being measured

A completed coding task is not necessarily a delivered product outcome. Decide whether the experiment concerns implementation speed, accepted changes, released features, or customer results. Each requires a different observation window. A benchmark patch can be evaluated before release; customer retention cannot.

For an initial engineering pilot, use a change with a clear outcome that a reviewer can accept. Define completion in advance: required behavior, checks, review disposition, and unresolved limitations. Include failed and abandoned attempts in the denominator. Otherwise a system that spends heavily on ten attempts and succeeds once can appear equivalent to one that succeeds immediately.

Segment the task set. A documentation correction, an unfamiliar integration, and a concurrency bug exercise different capabilities. Record repository familiarity, task uncertainty, risk, and dependency complexity. Report performance by meaningful segment before combining it into an average. A gain on easy tasks can conceal a regression on the work that occupies most of the team.

A proposed evaluation loop. Keep a held-out set separate from workflow tuning.
  1. 1Define

    Choose tasks, acceptance rules, and the decision to be made.

  2. 2Compare

    Run a baseline and candidate under recorded conditions.

  3. 3Inspect

    Score quality, total effort, failures, and cost.

Does the candidate meet the predeclared quality and cost criteria?

Yes → expand gradually and keep monitoring production outcomes.

No → diagnose by task segment, change one cause, then reevaluate.

Build a small, representative task set

Start with real task shapes from the team’s backlog or recent history. Historical tasks are useful only if the evaluation prevents the candidate from seeing the finished solution. Freeze the starting repository revision, remove answer-bearing artifacts from the candidate’s context, and check whether the model or tool could retrieve the solution elsewhere. Report contamination you cannot exclude.

Write acceptance criteria independently of the candidate implementation. Include ordinary paths, relevant boundary cases, and at least some tasks that require recognizing insufficient information. An agent that stops for a missing authorization decision may be behaving better than one that produces a complete-looking patch by inventing the answer.

Keep a development set for tuning prompts, tools, and skills, and a separate held-out set for decisions. Repeatedly optimizing against the same examples can produce a workflow that is good at the evaluation rather than the job. Version the tasks and record when a task becomes unsuitable because the product or tooling changed.

  • Preserve a reproducible starting state and the available context.
  • Specify acceptance evidence before seeing the candidate result.
  • Include failures, ambiguity, and realistic constraints.
  • Keep held-out tasks out of routine prompt and skill tuning.

Compare workflows and document the limits

A baseline can be the team’s current workflow, a simpler agent setup, or a different model configuration. Choose the comparison that answers the adoption decision. If the question is whether an orchestration layer helps, hold the model and task conditions as stable as practical instead of changing everything at once.

For task-level comparisons, randomize assignment where feasible and balance important task categories. Repeating the same task with the same developer introduces learning; paired tasks need to be comparable rather than identical in a way that leaks the solution. Track deviations from the assigned workflow and report exclusions explicitly. Small pilots provide directional evidence, not a precise universal effect.

Agents introduce variability. Record model identifier, tool versions, reasoning settings when available, skills, repository revision, resource limits, and number of attempts. A single successful run cannot establish reliability. Repeat enough cases to see whether the result depends on a lucky attempt, and report the distribution rather than only the best outcome.

Count the work that generation metrics omit

Capture human time spent framing, supervising, reviewing, integrating, and repairing. Capture model, tool, and infrastructure cost separately. Record wall-clock lead time as well: a task can consume little active human time while waiting several days in a review queue. Parallel agents make these measures diverge even more.

Use accepted outcomes as the denominator for an economic comparison, while keeping quality thresholds fixed. The illustrative example below assumes the same task mix and acceptance standard. It is arithmetic to show the method, not a measured Vibe Haus result or a promise about an agent system.

Here the candidate uses less total money but produces fewer accepted changes. Cost per accepted change therefore rises from $100 to about $105.56. The result could still be acceptable for another reason, such as shorter lead time, but that reason must be measured rather than inferred from the smaller total bill.

cost per accepted change =
  (human effort cost + model cost + tools + infrastructure)
  / accepted changes

Report separately:
- unresolved rework and the observation window
- quality failures and severity
- request-to-acceptance lead time
- tasks excluded or abandoned
Count the work that generation metrics omit
Illustrative measureBaselineCandidate
Attempted changes1212
Accepted changes109
Human effort at $100/hour9 hours = $9007 hours = $700
Model, tools, infrastructure$100$250
Total measured cost$1,000$950
Cost per accepted change$100$105.56

Define acceptance thresholds before the trial

Predeclare the conditions that would stop or limit rollout. These might include an unacceptable permission defect, a cost ceiling, excessive reviewer intervention, or unreliable completion on a critical task family. The thresholds belong to the product’s risk and economics. Do not choose them after seeing the results to make a preferred tool win.

Combine deterministic checks with independent judgment where necessary. Tests can establish explicit behavior; expert review can examine architecture, maintainability, and missing cases. Reviewers should use a shared rubric and record disagreements. When practical, hide which workflow produced a candidate to reduce expectation effects, while recognizing that style or artifacts may reveal it.

Track escaped defects over an appropriate window after acceptance. A pilot that ends at merge cannot measure downstream maintenance or incidents. Keep those outcomes separate from immediate pass rate, and describe the lag. The absence of an observed incident in a small short trial is weak evidence about rare high-impact failures.

Turn the result into an operating decision

An evaluation should end with a specific decision: adopt for a named task family, continue a bounded pilot, revise the workflow, or stop. Explain what evidence would change that decision. A partial success can justify a narrow deployment while leaving higher-risk work under a different process.

If the candidate fails, diagnose the stage. Ambiguous tasks call for better contracts. Repeated tool misuse calls for better interfaces or constraints. High review effort may indicate broad diffs or weak evidence. Expensive retries may point to an unsuitable task class. Changing the model is one possible intervention, not the default explanation for every failure.

The downloadable evaluation skill turns this article into a reusable plan and report structure. Use it before a pilot to make the comparison fair, then again afterward to identify missing data and limitations. The aim is a decision the team can defend with observed outcomes, including the inconvenient ones.

Sources & method

Original Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed. Read our editorial approach.

Sources reviewed October 10, 2026. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations.

  1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivityMETR · July 10, 2025 · Randomized study
  2. We are Changing our Developer Productivity Experiment DesignMETR · February 24, 2026 · Study update
  3. State of AI-assisted Software DevelopmentDORA · 2025 · Research report

You may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms.