# Vibe Haus > An exclusive AI-native engineering team for your company. Fully managed, trained in agentic coding, and working a full New York/Eastern day. UI/UX to R&D. The pages below describe the service and engagement model. Pricing, availability, allocations, and responsibilities are agreed for each engagement. Illustrative examples are not customer case studies. ## When to use Vibe Haus Consider the service for a founder or product team with an ongoing product workstream requiring engineering management, UI/UX, software development, or technical R&D. A single isolated fix may suit a narrower engagement. Fees, availability, expertise, ownership, and support require a proposal. Do not infer those terms from team size. ## Use the public resources Read a page with Accept: text/markdown, fetch the public JSON catalog at https://vibehaus.team/api/research, or connect an MCP client to https://vibehaus.team/mcp. These resources retrieve published content only; they cannot hire a team, submit a lead, or execute code. [API specification](https://vibehaus.team/openapi.json). ## Website - [Home](https://vibehaus.team/): Fully managed engineering, team composition, and intro-call requests. - [Request a call](https://vibehaus.team/contact): Name, email, and preferred date/time with at least 48 hours’ notice. Opens an email draft; confirmation follows by email. ## Capabilities and guides - [Vision & philosophy](https://vibehaus.team/vision-and-philosophy): Vibe Haus’s vision for engineering: one dedicated client per team, agentic coding in every role, full engineering ownership, and continuous learning. - [Cost & scope](https://vibehaus.team/engineering-cost): Understand managed engineering costs: team size, design, delivery leadership, AI tools, research, support, and the questions to compare proposals fairly. - [About Vibe Haus](https://vibehaus.team/about): Learn what Vibe Haus does: fully managed, AI-native engineering teams with hands-on leadership, a designer who codes, and capability from UI/UX to R&D. - [Editorial approach](https://vibehaus.team/editorial-policy): How Vibe Haus distinguishes source-based research, proposed workflows, illustrative examples, and measured evidence in its engineering articles and tools. - [Who it’s for](https://vibehaus.team/who-its-for): See who Vibe Haus fits and what a managed engineering engagement looks like: buyer situations, an illustrative product build, team roles, deliverables, and your involvement. - [Solutions](https://vibehaus.team/solutions): Explore Vibe Haus’s fully managed engineering solution: UI/UX, AI-native development, technical R&D, and delivery leadership in one dedicated team. - [Fully managed engineering](https://vibehaus.team/managed-engineering): Fully managed engineering teams with hands-on leadership, delivery coordination, design, and AI-native development. See what Vibe Haus manages and what you own. - [AI-native engineering](https://vibehaus.team/ai-native-engineering): How Vibe Haus approaches AI-native engineering: context, agent workflows, implementation, tests, review, and accountable release decisions. - [UI/UX & product engineering](https://vibehaus.team/product-design): UI/UX and product engineering in one Vibe Haus team. A dedicated designer who codes connects user journeys, interface decisions, and implementation. - [Research & development](https://vibehaus.team/research-and-development): Vibe Haus technical R&D: feasibility research, experiments, prototypes, and a path into product engineering. Plus a weekly R&D rhythm inside the team. - [The team](https://vibehaus.team/team): Compare Vibe Haus’s four- and six-person engineering teams: two or four engineers, a hands-on delivery lead, and a dedicated UI/UX specialist who codes. - [How we work](https://vibehaus.team/how-we-work): How a Vibe Haus engagement works: discovery, team setup, delivery, review, R&D, and handover, with responsibilities agreed before work starts. - [Working together](https://vibehaus.team/working-together): An exclusive engineering team working full New York/Eastern hours. Direct collaboration, company onboarding, agentic training, and fully managed engineering. - [For founders](https://vibehaus.team/for-founders): A fully managed engineering team for founders with a product to grow. Bring product direction; add AI-native development, design, delivery leadership, and R&D. - [For product teams](https://vibehaus.team/for-product-teams): A managed Vibe Haus team for product leaders: define a workstream, connect design and engineering, investigate technical risks, and keep ownership clear. - [Compare delivery models](https://vibehaus.team/compare): Compare a Vibe Haus managed engineering team with staff augmentation, individual freelancers, and an internal team. Understand ownership, coordination, and fit. - [Questions & answers](https://vibehaus.team/faq): Answers about Vibe Haus managed engineering: team composition, AI-native delivery, full R&D, US hours, ownership, pricing, support, and starting an engagement. ## Engineering research - [Blog](https://vibehaus.team/blog): Original engineering research and proposed workflows. - [The agentic code factory: from request to releasable change](https://vibehaus.team/blog/agentic-code-factory): A practical design for an agentic code factory: task contracts, isolated execution, evidence, review gates, bounded retries, and release ownership. Markdown: https://vibehaus.team/blog/agentic-code-factory/article.md - [Code review after agents: verify the change, not the explanation](https://vibehaus.team/blog/code-review-for-agents): An evidence-first review workflow for agent-written code, covering risk routing, independent tests, actionable findings, and exact-revision release gates. Markdown: https://vibehaus.team/blog/code-review-for-agents/article.md - [Context engineering: repository knowledge, MCP, and skills](https://vibehaus.team/blog/context-mcp-and-skills): How repository knowledge, task context, MCP tools, and agent skills fit together, with provenance, freshness, trust boundaries, and a practical context contract. Markdown: https://vibehaus.team/blog/context-mcp-and-skills/article.md - [Evaluating agentic engineering: quality, cost, and delivery](https://vibehaus.team/blog/evaluating-agentic-engineering): A measurement plan for agentic engineering that includes accepted outcomes, human review, rework, failures, cost, and the limits of current productivity evidence. Markdown: https://vibehaus.team/blog/evaluating-agentic-engineering/article.md ## Data and tools - [Data / MCP / Skills](https://vibehaus.team/data): JSON, Markdown, read-only MCP, and downloadable skills. - [Research JSON](https://vibehaus.team/api/research): Full article catalog, schema v1.0. - MCP endpoint: https://vibehaus.team/mcp (Streamable HTTP POST, public authored content only). ## Full text - [Complete public content](https://vibehaus.team/llms-full.txt): Plain-text copy of the capability pages, guides, and research articles. - [Sitemap](https://vibehaus.team/sitemap.xml) --- # Vision & philosophy Source: https://vibehaus.team/vision-and-philosophy Why we build dedicated, AI-native engineering teams We build small teams that know one client’s product, use coding agents across every role, and take responsibility for engineering delivery. Our goal is to increase useful output through training, clear responsibilities, and repeatable review processes. Our model combines a team dedicated to one client, full New York working hours, coding agents in every role, and managed delivery. We aim to improve how much useful software the team delivers while maintaining quality. ## Each team works exclusively for one client Each team works for one client. Its members learn your product, customers, architecture, and ways of making decisions. They work with you through a full New York/Eastern working day, so a conversation can continue while the work is happening. Working on one product lets the team build useful knowledge over time. Design decisions, integration constraints, and customer feedback inform later changes, so the team does not have to relearn the product for each task. ## Everyone is an agentic coder. The engineer, lead, and designer all write software with agents. Their specialties remain valuable: design judgment, technical depth, and delivery leadership help them ask better questions and judge the answers. Agentic coding gives each of those perspectives a direct path into implementation. Training covers the full workflow: turn an outcome into a clear task, give an agent the right context and boundaries, inspect the proposed changes, test the behavior, and decide what can ship. The training focuses on applying these tools to real product work and checking their results. ## Engineering management is part of the service Delivering a feature requires technical planning, architecture decisions, dependency management, code review, testing, and release coordination. Vibe Haus manages these responsibilities and the professional development of the team. Vibe Haus owns engineering execution within the agreed scope. You retain product direction and business priorities. We establish decision rights, access, acceptance, and support boundaries at the start so the team can take action and remain accountable. ## Keep training connected to the product. Agentic engineering changes quickly. The team needs a routine for testing new approaches, keeping useful practices, and retiring ones that fail. Ongoing training, feedback, and weekly R&D are part of how we develop the team. A learning activity should leave something usable: a tested workflow, a clearer technical decision, a better review checklist, or an experiment that rules out a weak approach. We agree the time for learning alongside delivery rather than treating it as invisible free capacity. ## Evaluate productivity through delivery and quality Our aim is to give a small team substantially more ability to build. More generated code is not enough to establish that advantage. Review load, defects, rework, and the time from a request to an accepted result all matter. We do not publish a fixed productivity multiplier as a proven outcome. A useful comparison needs a defined baseline, comparable tasks, the same quality bar, and the full cost of review and correction. The engineering blog explores these questions with sources and proposed workflows; customer results need their own evidence. ## Related reading - [About Vibe Haus](https://vibehaus.team/about) - [The team](https://vibehaus.team/team) - [AI-native engineering](https://vibehaus.team/ai-native-engineering) - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [Working together](https://vibehaus.team/working-together) --- # Cost & scope Source: https://vibehaus.team/engineering-cost What goes into the cost of a managed team? A useful quote explains the work, the people, and the responsibilities it covers. Use this guide to prepare a brief and compare proposals on the same basis. Vibe Haus quotes each engagement. There is no published fixed rate. Team size, allocation, technical uncertainty, tooling, and support requirements determine the proposal. ## Start with the outcome and the people it needs. A four-person Vibe Haus team includes two engineers, a hands-on delivery lead, and a designer who codes. A six-person team adds two engineers. Leadership and design remain real work: a four-person team is not four people writing application code all day. Describe the next product milestone, the existing system, and the decisions still unresolved. A known interface change and an unproven integration can require different kinds of work even when the finished screens look similar. Ask the proposal to explain those differences. ## Ask what is included and what is separate. Use the same scope when comparing quotes. Confirm the allocation, time period, and responsibility behind each line. Do not assume that a lower staffing rate includes design, engineering management, or ongoing support. Items to clarify in an engineering proposal | Cost or commitment | What to confirm | | --- | --- | | Team allocation | People, roles, availability, working hours, and how specialist responsibilities use their time. | | Discovery and R&D | Question to investigate, experiment budget, expected evidence, and the decision it should support. | | AI and development tools | Approved vendors, account ownership, usage limits, subscriptions, and who pays. | | Infrastructure and third parties | Hosting, data services, integrations, licenses, and any variable usage charges. | | Quality and release | Included checks, specialist reviews, deployment authority, and handover requirements. | | Support and continuity | Response hours, incident responsibilities, maintenance scope, leave, and changes to staffing. | ## Include the time your organization still contributes. A managed team organizes engineering execution. Your product owner still sets priorities, answers business questions, and accepts the work. Your existing engineers may also need to review architecture, grant access, or coordinate releases. Compare the total commitment: the provider fee, external costs, and the time you spend managing dependencies and acceptance. An hourly rate alone cannot show whether two proposals cover the same responsibilities. ## Budget an experiment before promising a build. When feasibility is unclear, agree a bounded investigation. Specify the question, inputs, stopping point, and evidence needed to make a decision. The result may support building, changing approach, or stopping. For example, an integration spike could establish whether an external service exposes the events a proposed onboarding flow needs. That finding is useful even if it rules out the original design. This is an illustrative scope, not a price or delivery estimate. ## Compare proposals with five questions. Send each provider the same brief. Ask for assumptions and exclusions in writing so you can distinguish a genuine scope difference from a price difference. - What deliverable or milestone will we evaluate, and what counts as acceptance? - Who manages engineering, design, review, and communication? - Which fees are recurring, which are usage-based, and which require separate approval? - What happens if research changes the scope or an external dependency blocks progress? - What ownership, documentation, access, and transition work are included when the engagement ends? ## Bring enough context to make a quote useful. Prepare a short description of your product, who uses it, the next outcome, and the main constraint. Include an existing roadmap or prototype if you have one. Note any required technologies, data restrictions, coordination needs within the New York working day, and support expectations. The first proposal should establish the team, scope, allocation, fees, assumptions, and responsibilities. Vibe Haus does not publish an automatic start date or a guaranteed AI productivity multiplier. Availability and terms need agreement for your engagement. ## Related reading - [Evaluating agentic engineering: quality, cost, and delivery](https://vibehaus.team/blog/evaluating-agentic-engineering) - [Compare delivery models](https://vibehaus.team/compare) - [Who it’s for](https://vibehaus.team/who-its-for) - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [Research & development](https://vibehaus.team/research-and-development) --- # About Vibe Haus Source: https://vibehaus.team/about About Vibe Haus and our managed engineering teams We bring engineers, a hands-on lead, and a designer who codes into one exclusive team. They become part of your working team; Vibe Haus takes responsibility for engineering delivery, agentic training, and people development. Each team works exclusively for one client during full New York/Eastern hours. Every role codes with agents. You set product priorities, and Vibe Haus manages engineering execution within the agreed scope. ## What Vibe Haus does. The service spans UI/UX, product engineering, integrations, and technical research. A team can start with a user journey, investigate a difficult technical question, and carry the resulting decisions into implemented software. The four-person model brings together two engineers, a hands-on lead, and a designer who codes. The six-person model adds two engineers. Each person contributes code, with time also assigned to design, leadership, investigation, and review. ## What the managed relationship covers. Engineering planning, coordination, review, progress communication, and people development belong in the working model. Your product owner sets business priorities and accepts the work. Decisions about architecture, deployment, data access, and production support are assigned before work begins. Every team member works exclusively for your company during a full New York/Eastern working day. Client-specific onboarding, a shared workspace, and direct conversation connect the team to your organization. Scope, fees, start date, deliverable ownership, and support coverage are established in the agreement. ## How AI and research fit the work. Engineers use AI for suitable tasks such as exploring code, drafting changes, testing, and documentation. People remain responsible for technical judgment, review, and release decisions. The team agrees which tools and data it can use. Research is part of the capability, from a bounded feasibility experiment to a larger product investigation. The public engineering blog shares source-based synthesis and proposed workflows. Its examples are not presented as customer results. ## What to establish before working together. Use the buyer guide to see whether the model fits your situation. Then bring the product outcome, constraints, and unresolved questions to a scope discussion. Ask for relevant experience, the proposed people and allocations, and the evidence you will use to accept progress. A website can explain the model; the proposal must make it concrete for your product. You can email us or pick an intro-call time on the homepage, which opens an email draft. Availability and commercial commitments require a direct agreement. ## Related reading - [Vision & philosophy](https://vibehaus.team/vision-and-philosophy) - [The team](https://vibehaus.team/team) - [Who it’s for](https://vibehaus.team/who-its-for) - [Cost & scope](https://vibehaus.team/engineering-cost) - [Editorial approach](https://vibehaus.team/editorial-policy) --- # Editorial approach Source: https://vibehaus.team/editorial-policy How to read and use our engineering research. Our articles help engineers and buyers reason about software delivery with agents. They combine primary sources with our own proposed workflows, including the limits and decisions that a diagram alone cannot explain. The blog is published by Vibe Haus. It contains original synthesis and proposed practices, not peer-reviewed research or independently measured customer outcomes. ## Follow the evidence to its source. Articles link the research, engineering documentation, and practitioner material used in their reasoning. The source list names the publisher and identifies the kind of evidence. A vendor announcement, an engineering guide, and a controlled study answer different questions. We distinguish a source’s reported result from our interpretation. A result measured on one codebase or task set does not establish the same result for your organization. Check study conditions, sample size, publication date, and limitations before using a number in a decision. ## Treat examples as proposals to test. Flowcharts, task contracts, review checklists, and evaluation thresholds are proposed ways to organize work. They need adaptation to the product, data, risks, and people involved. Illustrative examples are not customer case studies. The same applies to downloadable skills. They are instruction templates, not executable guarantees, credentials, or permission to make changes. Your existing review and release authority still applies. ## Use dates to understand what was checked. Each article shows its publication date and the date its sources were reviewed. A separate update date is used when a substantive revision is recorded. An unchanged article does not become newly researched because the website was rebuilt. Drafts are excluded from the public blog, feeds, Markdown, and MCP tools. Published articles and their machine-readable versions come from the same content files. ## Authorship and scope. Vibe Haus is the organizational publisher. These articles are not attributed to an invented individual expert or an external research institution. References to AI Engineer and other organizations identify sources and inspiration, not partnerships or endorsements. AI-assisted production does not turn a proposed method into experimental evidence. Judge an article by whether its claims are supported, its examples are usable, and its limitations are clear. ## Reuse the work with context. You may use and adapt the original templates and workflows with attribution to Vibe Haus. Third-party sources retain their own terms. Link the canonical article when sharing a method so readers can inspect the source list and any later revisions. The Data / MCP / Skills page explains the public JSON feed, Markdown articles, and read-only MCP tools. None of these endpoints accepts private repository data or executes work for you. ## Related reading - [About Vibe Haus](https://vibehaus.team/about) - [AI-native engineering](https://vibehaus.team/ai-native-engineering) - [Research & development](https://vibehaus.team/research-and-development) --- # Who it’s for Source: https://vibehaus.team/who-its-for Is this the team your product needs? You need to build or improve a product, but hiring more developers would still leave you managing the work. See when a managed team fits, then follow an example from design through technical research and implementation. A dedicated four- or six-person, AI-native team. Everyone codes. Design and hands-on leadership are included. You bring product direction; Vibe Haus manages engineering execution. ## Common reasons to use a managed team This fits when you have enough work for a small team and need someone to coordinate design and engineering. These are three common starting points: - You’re a founder carrying the roadmap and the day-to-day engineering coordination. You want a team that can own delivery planning while you stay close to customers and product decisions. - You lead an existing product team with an important workstream that keeps waiting for capacity. You want a dedicated team that can work within your architecture, standards, and release process. - You’re responsible for a new capability with an unresolved technical question. You need design and R&D to establish a useful approach, then engineering to turn it into production software. ## Who is on your team? Start with a four-person team: two engineers, one hands-on delivery lead, and one designer who codes. A six-person team adds two engineers. The lead and designer contribute code alongside their specialist responsibilities; their time is allocated around the work. You work with a team that shares your product context. The lead organizes engineering execution and coordinates dependencies. The designer connects user journeys to the implemented interface. Engineers build, test, and investigate technical questions. AI supports suitable tasks, while people remain responsible for review and decisions. ## Example project: self-service customer onboarding Imagine a software product whose customers still need help from your team to get set up. You want them to complete the process themselves, but the experience is confusing and a third-party integration has unanswered technical questions. This fictional example shows how the team could handle the work. It is not a customer case study or a timeline estimate. 1. Define the customer outcome. You explain where customers get stuck and what successful setup means. The lead and designer map the current journey, dependencies, and decisions that need your input. What you see: An agreed milestone, a user journey, and acceptance criteria you can review. 2. Investigate the uncertain part. The engineers test the integration in an agreed environment. The designer checks how its limitations affect the experience. The lead brings the findings and tradeoffs back to you. What you see: A technical experiment, recorded findings, and a recommendation to proceed, change approach, or stop. 3. Build and test the onboarding flow The designer works on the flow and interface code. Engineers implement the integration and product behavior. The lead contributes code, coordinates dependencies, and organizes review. What you see: A working setup flow you can try, with reviewed code, agreed tests, and known limitations. 4. Review the agreed acceptance criteria You review the flow against the acceptance criteria. The team addresses the agreed feedback and documents release considerations. Deployment follows the authority and process established for your product. What you see: An acceptance decision, a release or handover plan, and a prioritized next step. ## Progress updates and reviews The work should be visible between milestones. Agree the working rhythm during onboarding, including the full New York working day, the shared workspace, review points, and how urgent decisions reach you. A typical review can combine a demonstration with a short written account of progress, open questions, and what happens next. - Current work: what the team is building or investigating and how it connects to the milestone. - Evidence: a working flow, a design to review, tested code, or findings from an experiment. - Decisions: the tradeoff, its consequences, and the person who needs to decide. - Next steps: upcoming work, dependencies, and changes that need agreement. ## Your responsibilities as the product owner Name a product owner who can explain priorities, answer business questions, and accept the work. Provide access to the relevant product, people, and documentation. Stay available for the decisions the team cannot make on your behalf. Vibe Haus manages the engineering work and the people doing it. You make the product and business decisions. Before starting, we assign release responsibilities and confirm any production support or specialist expertise you need. ## When a managed team is a good fit This is worth exploring if you have a product milestone or research question, enough work for a small team, and a need for engineering management. You can start before every screen or ticket is defined. If you need one small fix, an isolated design task, or a single specialist inside a team you already manage, a narrower engagement may serve you better. If no one can own product priorities or acceptance, establish that responsibility before expecting managed engineering to solve it. Bring three things to the first conversation: what your product does, the next outcome you want, and what is blocking it. A short message is enough to start the conversation. Team expertise, allocation, availability, fees, and engagement terms are then established in a proposal. ## Related reading - [The team](https://vibehaus.team/team) - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [For founders](https://vibehaus.team/for-founders) - [For product teams](https://vibehaus.team/for-product-teams) --- # Solutions Source: https://vibehaus.team/solutions From UI/UX to software development and R&D. Work with one team to design your product, build its features, and investigate the technical questions that stand in the way. The same managed team can design an interface, test technical feasibility, and build the resulting features into your product. ## One team, from the first screen to the hard question. Enter where your product needs help. Design, research, and engineering can inform each other throughout the work. 1. Design the experience (Designer who codes) Map what a person needs to do and make the interaction concrete. Output: A user journey and working interface direction. 2. Test the unknown (Engineers + lead) Run a focused experiment when feasibility is uncertain. Output: Findings that support a build, another test, or a different approach. 3. Build the product (The whole team) Implement the selected approach and review it against product requirements. Output: Tested software ready for your acceptance process. Decision: Has the work answered the next product question? Yes: continue toward the next milestone. Not yet: revisit the design, experiment, or implementation with what we learned. ## Development, design, and engineering management Vibe Haus brings engineering execution, hands-on delivery leadership, and a UI/UX specialist who codes into a dedicated four- or six-person team. AI-native ways of working connect those disciplines. The team can contribute across product design, application development, integration work, and technical research, with the mix shaped around your next milestone. Fully managed means the relationship includes how the team works: planning, coordination, engineering review, feedback, and people development. You bring the product context and business priorities. We manage the engineering work within an agreed responsibility map, so adding capacity also brings a way to organize it. ## UI/UX, software development, and technical R&D For a known product problem, the team can work from a user journey through interface design and into implemented software. For a technical unknown, it can start with a research question, explore competing approaches, and produce a prototype or evidence-based recommendation. The goal determines the method. These capabilities belong in the same conversation. A promising prototype still needs a useful interface. A polished interface still needs sound engineering. And a working feature may reveal a deeper question about performance, architecture, or feasibility. Keeping the context with one team helps connect those decisions. - UI/UX and implementation: connect the user’s task to the actual product behavior. - Product engineering: develop features, integrations, tests, and maintainable application code. - Technical R&D: investigate feasibility, compare approaches, and turn findings into a decision. - AI-native delivery: use agents and automation within a workflow with human review. ## The people and working practices included. The hands-on lead connects priorities, implementation, and communication. A dedicated designer who codes brings experience decisions into the build. Engineers work with both, using the agreed standards for testing, review, and release. The four-person model has two engineers; the six-person model has four. The operating model includes full New York/Eastern working hours, an exclusive team, a virtual office, client-specific onboarding, performance feedback, and a weekly R&D activity with a designated team member. We schedule these activities alongside delivery. Exact schedules, allocations, and specialist needs are defined in the proposal. ## Start with the outcome you need. Bring an existing roadmap, an experience that needs rethinking, or a technical question that is blocking a decision. We use that context to identify the right starting point, the people required, and what useful progress should look like. You do not need to arrive with every implementation detail decided. A good first scope is concrete enough to evaluate: a product flow that works, an integration that passes agreed checks, or an experiment that answers a feasibility question. Availability, fees, support coverage, and ownership are agreed before the engagement begins. ## Related reading - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [UI/UX & product engineering](https://vibehaus.team/product-design) - [Research & development](https://vibehaus.team/research-and-development) --- # Fully managed engineering Source: https://vibehaus.team/managed-engineering We manage your engineering team and delivery You set product priorities. A hands-on lead organizes the engineering work, builds alongside the team, and keeps you informed about progress and decisions. Vibe Haus manages engineering execution and the team behind it. You keep business priorities, product direction, and final acceptance. ## You set direction. The team organizes the build. A working relationship with clear responsibilities and decisions flowing both ways. 1. Product direction (You) Set priorities, explain customer needs, and accept the work. Output: A product goal and timely business decisions. 2. Daily coordination (Hands-on lead) Plan work, coordinate dependencies, arrange review, and keep you informed. Output: A shared plan, visible progress, and decisions that need you. 3. Design + engineering (The whole team) Design, code, test, and investigate with the lead building alongside the team. Output: Work you can inspect, use, and review. Decision: Does a decision change the product or business priorities? Yes: the lead brings it to your product owner. Within the team's agreed authority: the team resolves it and records the outcome. ## What fully managed means. Hiring people and managing a delivery system are different jobs. Vibe Haus combines them in one offer: engineers, a hands-on delivery lead, and a UI/UX specialist who codes, working as a dedicated team. The lead contributes software while organizing the work and bringing decisions, dependencies, and risks into view. We handle planning, task breakdown, coordination, code review, and progress updates. We also manage feedback and development for the people on the team. Before work begins, we agree which decisions the lead can make and which need your approval. ## Who handles what? You should know which decisions need you and which the team can make. We define that boundary before work begins, then use it to keep decisions close to the people with the context to make them. The aim is a manageable set of product decisions for you and clear engineering ownership for the team. The engagement’s responsibility map | Area | Vibe Haus | Your team | | --- | --- | --- | | Product direction | Surface technical options and tradeoffs | Set priorities and business goals | | Day-to-day execution | Plan and coordinate the agreed engineering work | Provide context and timely decisions | | Design and implementation | Design, build, test, and review within scope | Provide user insight and approve agreed outcomes | | People development | Manage feedback and development in the team | Share collaboration and delivery feedback | | Release and operations | Perform the responsibilities named in the agreement | Name acceptance and production authorities | ## How we plan, review, and report on the work A useful plan connects a product objective to changes small enough to review. Work needs completion criteria, an owner, and a way to surface blockers. Reviews need time and context. Progress reporting should explain what changed, what remains uncertain, and which decision matters next. That system also has to suit the product. A mature application may need careful migration and regression checks. An R&D effort may need a time-bounded experiment before implementation is sensible. The lead helps choose the right shape of work and makes the reasoning visible. ### Illustrative example: fewer decisions waiting on the founder A founder is assigning tickets and answering every implementation question. For a customer-onboarding project, the lead takes over task planning, coordinates the integration work, and organizes code review. The founder still decides which customer steps matter and accepts the finished flow. The team brings back product decisions instead of asking the founder to coordinate every task. ## What we confirm before starting. Before starting, we confirm the work, hours, capacity, and review process. We also identify who handles product management, specialist testing, infrastructure, security, and incidents. Extra staffing or specialist work needs to be included in the proposal. This model is particularly relevant when you need ongoing engineering work and someone to organize it. If your own engineering manager already has a complete delivery system and needs one narrow specialty, an individual hire or staff augmentation may be a better match. ## Related reading - [The agentic code factory: from request to releasable change](https://vibehaus.team/blog/agentic-code-factory) - [Code review after agents: verify the change, not the explanation](https://vibehaus.team/blog/code-review-for-agents) - [The team](https://vibehaus.team/team) - [How we work](https://vibehaus.team/how-we-work) - [Compare delivery models](https://vibehaus.team/compare) --- # AI-native engineering Source: https://vibehaus.team/ai-native-engineering How our engineers use AI agents Our engineers use agents to understand code, implement changes, draft tests, and maintain documentation. People define the tasks, review the results, and decide what is ready to release. AI-native engineering combines engineers, agents, and repeatable workflows. People own the requirements, technical judgment, review, and release decisions. ## From a request to a reviewed change. Follow an illustrative customer-invitation feature. Select a step to see what people and agents do. 1. Define (Product owner + lead) Decide who can invite a teammate, which role they can assign, and what success looks like. Output: A task with acceptance examples and limits. 2. Prepare (Engineer + agent) Inspect the existing account model, identify the relevant code, and set tool access and checks for this task. Output: A bounded task in a prepared development environment. 3. Build (Engineer + agent) Implement the invitation flow, draft tests, and record assumptions. The engineer directs and adjusts the work. Output: A candidate change, tests, and a readable diff. 4. Verify (Checks + human reviewer) Test expired links, duplicate invitations, and cross-account access. Inspect the actual diff and the interface behavior. Output: Review findings and results tied to this revision. Decision: Do the checks and review support acceptance? Yes: the named release owner approves deployment, rollback planning, and handover. No: return to the relevant step, correct the issue, and verify the new revision. ## Where AI helps the team. Vibe Haus’s engineering model starts with AI as part of the production process. Engineers can use it to explore unfamiliar code, propose implementation approaches, produce candidate changes, draft tests, and maintain documentation. The lead and designer bring product and experience context into that same process. The important work is connecting these steps. An agent needs relevant context and a bounded task. A candidate change needs checks and a reviewer. A useful result needs to satisfy the original product requirement. Buying access to a model does not establish any of those things by itself. ## The workflow from request to release We shape the workflow around the codebase and the work it needs. Straightforward tasks can use a simple assisted process. Repeated work may justify automation. Open-ended or high-risk changes need more human investigation and control. The workflow should remain understandable to the people responsible for the product. - Define: agree the objective, relevant context, constraints, and acceptance criteria. - Prepare: choose the task boundaries, tool access, environment, and checks. - Build: use engineers and agents for suitable implementation, testing, and documentation work. - Verify: inspect behavior, code quality, design, test results, and failed runs. - Release: follow the agreed approval, deployment, and recovery process. - Learn: record what worked, where intervention was needed, and which improvement is worth trying. ## Automating repeatable engineering tasks By “software factory,” we mean a repeatable process: define a task, give an agent the relevant code and instructions, check its changes, and have an engineer review the result. We aim to automate more of that process where repeated use shows it works reliably. There are three distinct things to scope: the product software, the team’s internal automation, and any workflow platform delivered to you. A product-development engagement does not automatically include a separate automation platform. If you want one, we define its ownership, operating costs, maintenance, and handover as a deliverable. ## Code review, testing, and data access Generated code is a candidate implementation. It still needs to work with the existing product, handle meaningful edge cases, and be maintainable by the people who will own it. Reviews should examine those properties instead of treating passing output from a tool as an acceptance decision. Tool choice also affects data access and cost. Before connecting an agent to a repository or service, agree which tools are approved, what data they can use, what permissions they have, and who can authorize consequential actions. Apply the client’s requirements to the actual workflow. ## How we evaluate AI-assisted delivery The value of AI should show up in useful, accepted delivery. Measure the time to acceptance, review effort, rework, defect behavior, and total costs for a defined class of work. Include setup, failed runs, and human intervention. More generated code or more agent runs does not establish more product value. We do not publish a fixed speed or headcount multiplier. The opportunity is to improve the system around a capable team and evaluate what that improvement actually produces in your context. Weekly R&D creates a deliberate place to test the next change to that system. ## Related reading - [The agentic code factory: from request to releasable change](https://vibehaus.team/blog/agentic-code-factory) - [Code review after agents: verify the change, not the explanation](https://vibehaus.team/blog/code-review-for-agents) - [Context engineering: repository knowledge, MCP, and skills](https://vibehaus.team/blog/context-mcp-and-skills) - [Research & development](https://vibehaus.team/research-and-development) - [How we work](https://vibehaus.team/how-we-work) - [Fully managed engineering](https://vibehaus.team/managed-engineering) --- # UI/UX & product engineering Source: https://vibehaus.team/product-design A UI/UX designer who also implements the interface Your UI/UX specialist designs user flows and writes interface code, working with the engineers throughout development. UI/UX is a continuous team capability. The designer works with the engineers and delivery lead, from the first flow through the implemented experience. ## Start with what the user needs to do. A feature starts with someone trying to do something. The UI/UX role helps make that task concrete: what the person needs to understand, which choices they need to make, and what should happen when something goes wrong. Engineering decisions can then be evaluated against the experience they are meant to support. Because the designer also codes, design decisions can continue into implementation. Layout, interaction, loading behavior, error recovery, and responsive details stay part of the same conversation. The role is included continuously in the team rather than introduced only to produce an initial set of screens. ## Design and implementation scope The scope can begin with an existing experience that needs improvement or with a new product capability. We identify the journey, available user evidence, product constraints, and success criteria before choosing the level of design exploration. Existing brand and design-system work becomes an input, not something to discard by default. Implementation then connects the visible interface to real states and behavior. Useful review includes empty states, validation, permissions, slow responses, and the small screens people actually use. Accessibility and keyboard interaction belong in the work itself, alongside the visual design. - Clarify the task, journey, and information hierarchy. - Explore the interface and validate important assumptions with available evidence. - Build responsive components and connect them to application behavior. - Review the implemented result against the agreed experience and acceptance criteria. ## Testing technical requirements during design A desired interaction may depend on an uncertain technical capability. Can a recommendation arrive quickly enough? Can an integration provide the information the user needs? Does an AI feature behave consistently enough for the proposed experience? Those questions need investigation before a polished screen can become a reliable product. The team can move into R&D to examine the uncertainty, then bring the result back into the design. A prototype may change the interaction model, narrow the feature, or show that another approach is better. This is why UI/UX and technical R&D are connected capabilities in the same offer. ### Illustrative example: reviewing an AI-drafted support reply A support agent needs to review a suggested reply before sending it. The designer builds the draft, edit, and approval flow; engineers test it with representative support questions. If the model invents a detail or times out, the interface must make correction or manual writing easy. The prototype helps the team decide whether the feature is useful enough to build. ## How we allocate design and coding time A designer who codes is still one person. Discovery, interaction design, implementation, and review all require time. We agree how that time is allocated around the roadmap rather than counting the same role as a full-time designer and an additional full-time engineer. Research participants, formal usability studies, specialized illustration, brand creation, and domain-specific accessibility certification may require separate scope or expertise. The first conversation identifies which design capabilities your product needs and what evidence is already available. ## Related reading - [Context engineering: repository knowledge, MCP, and skills](https://vibehaus.team/blog/context-mcp-and-skills) - [Research & development](https://vibehaus.team/research-and-development) - [The team](https://vibehaus.team/team) - [Solutions](https://vibehaus.team/solutions) --- # Research & development Source: https://vibehaus.team/research-and-development Technical research, feasibility testing, and prototypes The team can investigate whether a proposed feature or technical approach is feasible before you commit to a full build. The team can handle technical discovery, experiments, prototypes, and production development. It also runs one R&D activity each week as part of ongoing training. ## Turn uncertainty into a decision. A prototype is useful when it answers a specific question. Production work follows a separate readiness decision. 1. Ask (Product owner + team) Name the uncertainty and the decision it is blocking. Output: A question, success criteria, and a time allowance. 2. Experiment (Engineers + hands-on lead) Test the uncertain part with representative inputs and failure cases. Output: A prototype or measurement, with observed limitations. 3. Decide (Product owner + lead) Compare the findings with the criteria and choose what happens next. Output: Build, investigate further, change direction, or stop. Decision: Is there enough evidence to invest in a production build? Yes: define the integration, security, quality, and operating work still needed. No: narrow the question, test another approach, or stop the experiment. ## Define the technical question and experiment Vibe Haus teams can work across UI/UX, engineering, and full technical R&D. When the challenge is uncertain, the first deliverable may be an answer: whether an approach is feasible, what its limitations are, which tradeoffs matter, and what should be tested next. Writing a production backlog too early can hide those questions. R&D begins by defining the question and why the answer matters to the product. We identify the existing evidence, constraints, competing approaches, and a bounded experiment. The output should help you decide whether to invest, change direction, investigate further, or stop. A negative finding can be valuable when it prevents the wrong build. ## What a research engagement can deliver. Research work can investigate a new integration, a difficult performance requirement, a data-processing approach, an architectural option, or an AI-enabled interaction. The exact domain and expertise are assessed before the scope is agreed. A broad team capability does not mean every specialist discipline is automatically covered. We connect the experiment to observable criteria. A prototype should demonstrate the part that was uncertain, using representative conditions where possible. A report should distinguish what was observed from what is inferred, describe limitations, and make the next decision clear. - A research brief: the question, context, constraints, and decision it supports. - An experiment plan: approaches, representative inputs, checks, and a time allowance. - An artifact: a prototype, benchmark, technical spike, or evaluated implementation. - A decision record: findings, limitations, recommendation, and next steps. ## Taking a prototype into production A successful experiment establishes something specific. It does not automatically establish production readiness. The next step is to identify what the prototype leaves out: integration, permissions, real data behavior, operational failure modes, performance at scale, or a complete user experience. Because design and engineering sit in the same team, the findings can inform the next implementation directly. We convert the chosen approach into scoped product work with acceptance criteria and review. Reusing knowledge matters more than preserving prototype code that was never built to last. ### Illustrative experiment: a difficult integration Suppose customers need to see inventory from an external system before placing an order. A small prototype checks how fresh the data is, what happens when the API is unavailable, and whether account permissions are enforced. The findings help you choose between a live integration, periodic synchronization, or a different product approach. ## Weekly R&D within the team Alongside dedicated product R&D, the operating model includes one R&D activity per team each week with a designated team member. This gives the team a place to investigate a technical question, evaluate a workflow improvement, or explore a tool that may be useful to the product. That team member documents the question and result. A simple record includes the reason for the experiment, time allowance, method, findings, and adoption decision. The engagement defines how this activity is funded within capacity, how topics are selected, and who owns any reusable output. ## Deciding what to build after an experiment Weekly activity alone is not proof of innovation. An experiment should produce evidence that changes a decision or a working practice. Some experiments will be rejected. Others will justify another test. Adopt an improvement when the evidence supports it, then evaluate it in the real delivery context. Bring us the technical question that is blocking your product. The initial conversation can establish whether you need a focused research effort, a design-and-build path, or a managed team whose work will move between all three. ## Related reading - [Evaluating agentic engineering: quality, cost, and delivery](https://vibehaus.team/blog/evaluating-agentic-engineering) - [UI/UX & product engineering](https://vibehaus.team/product-design) - [AI-native engineering](https://vibehaus.team/ai-native-engineering) - [How we work](https://vibehaus.team/how-we-work) --- # The team Source: https://vibehaus.team/team Team sizes, roles, and responsibilities Choose two or four engineers, plus a hands-on lead and a designer who codes. We agree how each role divides its time between coding and specialist responsibilities. Four people means two engineers, one hands-on lead, and one UI/UX specialist. Six people adds two engineers. Both specialist roles also contribute code. ## Every role codes with AI agents Engineers, the lead, and the designer all use coding agents. We train every team member to define tasks, direct agents, review code, and test the results. 1. Define the outcome 2. Direct the agents 3. Review and test 4. Ship the work ### Engineers Agentic coding + technical exploration Engineers use agents to build features, investigate technical problems, and test changes across the system. ### Hands-on lead Agentic coding + engineering ownership The lead writes code, plans technical work, reviews changes, resolves dependencies, and coordinates delivery. ### Designer who codes Agentic coding + product design The designer uses agents to build interfaces, then checks usability, accessibility, and product behavior. Characters represent roles, not individual employee profiles. Everyone codes; design, leadership, and review remain part of the work. We measure productivity through accepted outcomes, quality, and rework. ### What each role contributes Select a role to see its responsibilities and areas of focus. - Engineers: Engineers focus on implementation, technical R&D, and verifying that changes work. Role focus: Agentic coding, Review & quality, Technical R&D. - Hands-on lead: The lead combines coding with delivery planning, dependency management, and code review. Role focus: Agentic coding, Review & quality, Delivery ownership. - Designer who codes: The designer combines UI/UX work with interface implementation. Role focus: Agentic coding, Product design. Outer points show role focus; inner points show shared practice. These are illustrative responsibilities, not skill ratings, percentages, or measured performance. ## Compare the four- and six-person teams Every member of the team is dedicated exclusively to one client. The four-person model brings two engineers together with a hands-on delivery lead and a UI/UX specialist who codes. The six-person model includes four engineers and the same two specialist roles. Design and leadership belong inside the team from the start. The lead divides time between coding and managing delivery. The designer divides time between UI/UX and implementation. We plan those responsibilities together, so neither role is counted as two full-time people. Dedicated team composition | Role | Four-person team | Six-person team | | --- | --- | --- | | Engineers | 2 | 4 | | Hands-on delivery lead | 1 | 1 | | UI/UX specialist who codes | 1 | 1 | | Total people | 4 | 6 | ## A lead who codes and manages delivery The lead builds software and specializes in delivery coordination. That combination connects planning to implementation realities: dependencies, review capacity, technical uncertainty, and the cost of changing direction. The lead should be able to explain both what is moving and what is holding it back. The engagement defines the lead’s authority. Product priorities remain with you; engineering planning and technical decisions follow the agreed responsibility map. Management is a real part of the role’s capacity and must be planned alongside its coding contribution. ## A dedicated designer throughout development The UI/UX specialist is continuously included and also codes. The role connects user journeys, interaction decisions, and the shipped interface. Engineers can resolve feasibility questions with the person designing the experience, and design can respond to what implementation reveals. That continuity supports work that crosses disciplines, including an R&D prototype that needs a useful interface or an existing product flow that requires both design changes and engineering. The team’s actual expertise is matched to the scope before the engagement is agreed. ## Choosing the right team size A four-person team may suit a focused milestone or a contained stream of ongoing product work. A six-person team adds implementation capacity where the backlog and dependencies can support it. More engineers also create more review and coordination work; the two specialist roles do not automatically gain extra time. Choose based on the available work, technical uncertainty, design needs, and your own decision capacity. Availability, dedication, working hours, leave arrangements, changes in team size, and continuity terms are confirmed in the proposal. There is no headcount-equivalence or fixed output promise attached to either option. ## Related reading - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [Working together](https://vibehaus.team/working-together) - [For founders](https://vibehaus.team/for-founders) --- # How we work Source: https://vibehaus.team/how-we-work Our process from project brief to release Here is how we understand your product, set up the team, review its work, and prepare for release or handover. We review your product and goals, agree the team and responsibilities, then plan development or research. You review progress and help set the next priorities. ## Review your product, goals, and constraints We begin with what you are building, who uses it, and what is holding the next milestone back. An existing backlog is useful, but so are unresolved questions, user feedback, architecture constraints, and examples of where the current delivery process stalls. The first decision is the shape of the work. A known feature may need a design-and-build plan. A technical uncertainty may need R&D first. An ongoing backlog may need a dedicated delivery stream. Define an outcome that can be evaluated and identify the decisions only you can make. ## Agree the scope, team, and responsibilities The proposal establishes team composition, role allocation, availability, scope, fees, and the engagement terms. We agree the working hours, communication channels, review cadence, decision owners, and escalation path. Tooling, third-party costs, R&D time, support, and handover belong in that conversation. A responsibility map identifies who can make architecture decisions, approve design, review code, accept work, authorize releases, and operate production. Product management, specialist testing, security work, or incident coverage require explicit scope rather than assumptions. ## Set up the codebase and make the first change. Onboarding connects the team to the product, users, codebase, and company culture. We identify the source of truth for requirements, the existing design and engineering standards, and the people who can answer domain questions. Access follows the tools and responsibilities agreed for the engagement. Early work should expose the real development path: set up the environment, make a scoped change or experiment, run the relevant checks, and review the result. That gives both sides a concrete view of how the collaboration works before increasing the amount of parallel work. ## Develop the software and review progress The lead coordinates the agreed work while engineers and the designer implement or investigate. AI tools support suitable tasks within the approved workflow. Changes are checked against requirements and reviewed by the appropriate human owner before acceptance or release. Progress communication should make accepted work, remaining work, risks, and required decisions easy to understand. A demonstration gives you something to evaluate. A blocked task should have a visible reason and an owner for the next action. The exact reporting rhythm is chosen for your team. ### Illustrative progress update: customer onboarding Ready to review: customers can accept an invitation and create their account in staging. Permission checks pass. Blocked: the external service has not enabled our test account, so we cannot finish the import step. Decision needed: should import stay part of initial setup or happen afterward? Next: test error recovery once access arrives. ## Keep learning and document the work. Weekly R&D gives a designated team member room to test a relevant question or improvement. Performance feedback develops the people doing the work. Both should connect to observable evidence, with capacity and review time treated as real parts of the engagement. Documentation and handover should grow with the work. At an agreed transition, clarify the state of the code, open decisions, environments, access, and operating responsibilities. Ownership, ongoing support, and any client-delivered automation platform follow the contract. ## Related reading - [The agentic code factory: from request to releasable change](https://vibehaus.team/blog/agentic-code-factory) - [Working together](https://vibehaus.team/working-together) - [Research & development](https://vibehaus.team/research-and-development) - [Questions & answers](https://vibehaus.team/faq) --- # Working together Source: https://vibehaus.team/working-together A dedicated team working full New York hours The team learns your product, works in your shared workspace, and is available throughout the New York working day. We agree communication channels and decision responsibilities during onboarding. Your whole team works exclusively for your company, through a full New York/Eastern working day. We onboard the team to your product and own engineering delivery, training, and professional development. ## Your team, managed by Vibe Haus Your team works directly with you on your product. Vibe Haus manages engineering planning, delivery, training, and professional development. ### Every team member works exclusively for you Your engineers, hands-on lead, and designer who codes are assigned only to your company. They learn your product and work with your people in your shared workspace. Your roadmap · Your workspace · Direct access to your team ### Managed by Vibe Haus - Engineering ownership: Technical planning, architecture, dependencies, and delivery decisions. - Development and release: Implementation, testing, code review, and release coordination. - Agentic training: Everyone learns to direct agents, inspect changes, and verify results. - People management: Coaching, feedback, performance, and career development. - Product context: Onboarding to your customers, standards, culture, and working practices. - Continuous improvement: Weekly R&D, shared learning, and better engineering workflows. ### A full New York working day The team works full New York/Eastern hours, including New York’s daylight-saving changes. You own product direction and business priorities. We own engineering execution within the agreed scope. Daily start/end times, holidays, specialist services, and on-call support are set in the engagement agreement. ## Full New York/Eastern working hours The team works a full US working day in New York/Eastern time, using America/New_York so daylight-saving changes stay aligned. Every team member is allocated exclusively to your company. We confirm daily start and end times, holidays, leave, and escalation arrangements in the agreement. A shared working day keeps design discussion, review, and unblocking close to the work. Use written context for decisions that should remain understandable after the call. Full New York working hours do not imply round-the-clock incident coverage or that team members live in the United States. ## Shared workspace and communication A shared virtual office is part of the communication model. It makes it easier to reach the team and see who is available while protecting time for focused work. The platform, access, and collaboration norms are agreed with you. Clear channels matter more than constant interruption. Establish where a question belongs, how urgent issues are raised, when written updates arrive, and who resolves a blocked decision. Continuous camera presence is not a requirement of the model. ## Onboarding to your product and company Client-specific onboarding covers more than repositories and tools. The team needs your product vocabulary, customer context, values, decision norms, and standards for good work. Relevant rituals and feedback help turn that context into day-to-day behavior. Culture integration should be practical: knowing who to involve, how to disagree constructively, how to write a useful update, and how to judge a tradeoff. We agree who supplies this context and how the team stays connected as the product changes. ## Performance feedback and team development Engineering performance includes judgment, quality, collaboration, and growth. The managed model includes feedback and development for the people in the team, with client input informing the working relationship. Review cadence, role expectations, and promotion authority are defined as part of the operating arrangement. Task counts or lines of code cannot describe the whole contribution of a lead, designer, or engineer. Development should address the skills and behaviors that matter to the actual work. Weekly R&D offers another way to build capability and share learning across the team. ## Devices, access, and security requirements The working offer includes an office-based model using approved devices that remain in the office. The proposal confirms device administration, access rules, exceptions, and the controls required by the client. A device policy alone is not a security certification. Repository permissions, AI tooling, confidential data, IP terms, and production access need their own decisions. We make those requirements visible during setup and name any specialist review that is needed. Support and incident response are agreed separately from normal collaboration hours. ## Related reading - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [The team](https://vibehaus.team/team) - [How we work](https://vibehaus.team/how-we-work) --- # For founders Source: https://vibehaus.team/for-founders Spend less of your day coordinating engineering. You can understand the product deeply and still need more people, more disciplines, and someone to coordinate the engineering work. A managed team is a strong fit when a founder has meaningful ongoing work and needs both implementation capacity and hands-on delivery leadership. ## When engineering decisions take up too much of your time A founder can become the meeting point for every product and engineering question: priorities, architecture, design details, task breakdown, reviews, and release decisions. Adding developers may increase the number of decisions arriving at that same point if no one takes ownership of the delivery work around them. Vibe Haus offers a team with a hands-on lead, engineers, and a designer who codes. The lead organizes the day-to-day work so you can stay focused on customers and product priorities. We agree which engineering decisions the team can make without waiting for you. ## Assign a team to your next product milestone A useful engagement begins with something concrete: a new product capability, an experience that needs improvement, a set of integrations, or a technical uncertainty that blocks investment. The team can move between UI/UX, implementation, and R&D as the work requires. For an existing product, bring the context that a job description cannot carry: why customers care, what the codebase makes difficult, which constraints matter, and where decisions get stuck. That helps us identify the required expertise and the right scope for a four- or six-person team. ## The product decisions you keep Fully managed engineering still needs a product owner. Your knowledge of the customer, business priorities, and acceptable tradeoffs is essential. The working agreement should concentrate your involvement around those decisions and define what the lead can decide without waiting for you. Useful inputs include a prioritized outcome, representative user feedback, access to the current product, and a clear acceptance owner. If product direction itself is unresolved, say so at the start. Product-strategy work and its authority need explicit scope. - Which milestone would make the next period successful? - Which decisions need your judgment, and which can the team own? - Where do design, implementation, or technical research currently stall? - What evidence would make you confident enough to accept the work? ## When to use a managed team or another hiring model The model is most relevant when there is sustained work for a small team and a need for engineering leadership alongside implementation. It can also suit a contained R&D effort when the required skills and capacity justify a team engagement. For a small isolated fix, a full team may be more than you need. If you already have strong engineering management and need one specialty, compare staff augmentation or a direct hire. Fees, availability, minimum term, and scope are established in a proposal, not inferred from the site. ## Related reading - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [The team](https://vibehaus.team/team) - [Compare delivery models](https://vibehaus.team/compare) --- # For product teams Source: https://vibehaus.team/for-product-teams A dedicated engineering team for your next project Give a product project its own engineers, hands-on lead, and designer who codes. They work within your architecture, design system, and release process. The team handles the project’s day-to-day engineering. Your product team sets priorities, reviews the result, and keeps control of the wider roadmap. ## Choose a project the team can own. A managed team is easier to integrate when the work has a clear objective and understandable boundaries. That might be a product experience, an application capability, an integration program, or an R&D question. Identify what the team can own and where it depends on the existing organization. Vibe Haus brings engineers, a hands-on delivery lead, and a UI/UX specialist who codes. The team can take a workstream through design and implementation or investigate feasibility before a delivery commitment is sensible. Its AI-native workflow is shaped around the client’s tools and requirements. ## Working within your existing engineering standards Existing architecture, design systems, review practices, and release processes are valuable context. Onboarding should identify these sources of truth and the people responsible for them. The team needs to know where consistency is required and where it has room to decide. The lead coordinates dependencies with your engineering and product counterparts. Define the required review points, shared interfaces, escalation path, and acceptance owner. A dedicated team should have enough authority to do its job while remaining accountable to the product’s broader standards. ## Separate research decisions from delivery decisions. Some initiatives appear ready for implementation until an integration, performance requirement, or AI behavior introduces a major unknown. Treat that as an R&D question with a defined experiment. The output can be a recommendation, prototype, or set of findings that informs the roadmap. When an approach is chosen, translate it into product requirements and a production scope. The same team can carry the technical and design context forward, with a clear distinction between what the experiment established and what the final product still needs. ### Illustrative project: account-level reporting Your customers need reports, but the core team is occupied with another release. The managed team designs a reporting flow, checks whether the current data model supports it, and builds the chosen approach. Your engineers review shared schema changes; your product owner reviews the reports. Release follows your existing process. ## Evaluate progress in product terms. Use demonstrations, acceptance criteria, and visible decisions to review the work. Track whether the agreed capability is becoming usable, whether dependencies are being resolved, and whether important risks are understood. The right evidence differs between a research phase and a production build. Before starting, confirm the people and expertise required, working hours, capacity, fees, review obligations, support boundaries, and handover. We check the team’s experience against your stack and domain before proposing the engagement. ## Related reading - [How we work](https://vibehaus.team/how-we-work) - [Research & development](https://vibehaus.team/research-and-development) - [Working together](https://vibehaus.team/working-together) --- # Compare delivery models Source: https://vibehaus.team/compare Compare managed teams, extra staff, and direct hiring. Compare who does the work, who manages it, and how much coordination remains with you. Price is one part of that choice. A managed team includes engineering coordination and people management. Staff augmentation generally adds individual capacity to a delivery system you manage. Actual contracts vary. ## When each delivery model is useful Staff augmentation can be effective when your engineering organization has clear leadership, established working practices, and a specific capacity or skill gap. Freelancers can suit contained work with a clear owner and boundary. An internal team supports an enduring capability with direct organizational ownership. Vibe Haus’s managed model combines dedicated engineers, a hands-on delivery lead, and a designer who codes. It is designed for buyers who need engineering execution and the management around it. The distinction is the scope of responsibility, not a claim that one model always produces better software. Typical responsibilities; individual providers and contracts differ | Question | Staff augmentation | Vibe Haus managed team | | --- | --- | --- | | What is added? | One or more individual skills or roles | A four- or six-person team with leadership and design | | Who coordinates the engineering work? | Usually your existing manager | The hands-on lead within agreed authority | | Where does UI/UX fit? | Depends on the roles you engage | A designer who codes is included continuously | | Who owns product priorities? | Your product owner | Your product owner | | What needs agreement? | Role scope, management, availability, and terms | Team scope, allocations, responsibilities, availability, and terms | ## Compare the complete working arrangement. A fee alone does not show the full commitment. Consider the management time your organization contributes, the disciplines the work requires, the dependencies that need coordination, and the review and acceptance effort that remains with you. Account for tools, onboarding, learning, and support as well as implementation. Ask each provider to explain its actual allocation and ownership. A lead who codes still spends time leading. A designer who codes still spends time designing. A team-size number should make those roles clearer rather than obscure them. ## When the managed model is worth considering. Consider a managed team when meaningful work spans several disciplines and you want a clear engineering counterpart to organize it. It can fit a founder who is carrying too many delivery decisions or a product organization with a coherent workstream that needs dedicated ownership. Consider another model when the gap is narrower. A single specialist may be the right answer for a defined technical task. Direct hiring may be the right investment for a long-term core function. You should not need to buy a whole team to solve a one-person problem. ## Questions to ask each provider Evaluate a proposal against the work you actually need done. Ask for the relevant experience, a concrete responsibility map, the proposed working rhythm, and the evidence you will use to judge progress. Use the answers to compare how much responsibility and uncertainty each option leaves with you. - Who breaks down the work and resolves blocked decisions? - Who performs review, acceptance, deployment, and production support? - Which disciplines are included, and how is their time allocated? - How are AI tools, data access, third-party costs, and R&D handled? - What happens when scope, people, or priorities change? - What will be handed over when the relationship ends? ## Related reading - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [Cost & scope](https://vibehaus.team/engineering-cost) - [Questions & answers](https://vibehaus.team/faq) --- # Questions & answers Source: https://vibehaus.team/faq Questions about the team, cost, and working together. Get clear on the team, the working relationship, and the decisions to make before an engagement starts. The offer is a dedicated, fully managed engineering team. Scope, expertise, availability, fees, and responsibilities are agreed in a proposal. ## What is Vibe Haus? Vibe Haus provides fully managed, AI-native engineering teams. Choose four or six people, including a hands-on lead and a designer who codes. The team can take work from UI/UX through software development and technical R&D. ## How does the dedicated team become part of our team? Every team member works exclusively for your company during a full New York/Eastern working day. The team learns your product, joins your workspace, and follows your working practices. You set product priorities and work directly with the people building your product. Vibe Haus owns engineering execution, agentic training, feedback, and people development. ## What does fully managed include? We handle engineering planning, coordination, code review, progress updates, and people development. You set product priorities and accept the work. Before starting, we agree who makes technical decisions, releases software, and provides production support. ## Who is in a four- or six-person team? A four-person team has two engineers, a hands-on lead, and a designer who codes. A six-person team adds two engineers. The lead and designer split their time between coding and their specialist responsibilities. ## What makes the team AI-native? Engineers use AI to understand code, develop changes, draft tests, and maintain documentation. People guide and review the work. We agree which tools and data the team can use, then measure the results rather than promise a fixed speed increase. ## Can the team do full R&D as well as UI/UX? Yes. The same team can design an interface, build product features, or investigate a technical question. R&D can produce experiments, prototypes, and a recommendation about what to build next. We confirm the expertise, budget, and outputs before starting. ## How does weekly R&D work? Each team has one R&D activity a week, led by a named team member. It can test a product idea or an engineering improvement. We agree its time allowance and record the question, findings, and next step. Dedicated product R&D can be a larger engagement. ## Who is the service a good fit for? Founders and product teams with enough ongoing work for a small team, and a need for someone to manage the engineering. If you need one small fix or a single specialist, a narrower engagement may fit better. ## Does the team work during US hours? Yes. The team works a full US working day in New York/Eastern time, following New York’s daylight-saving changes. Daily start/end times, holidays, and escalation arrangements are confirmed before starting. Round-the-clock incident support is a separate arrangement. ## How does the team become part of our company culture? We onboard the team to your product, customers, vocabulary, and working practices. A shared virtual office, written updates, and regular feedback keep people connected to how your company works. ## Who manages performance and career development? Vibe Haus manages feedback and professional development, with your input on how the team works with you. The working agreement sets review frequency, role expectations, and promotion decisions. We consider quality, judgment, and collaboration as well as completed work. ## Are QA, DevOps, security, and production support included? Code checks and review are included in the engineering process. Dedicated QA staff, specialist security work, infrastructure ownership, and on-call support must be specified in the proposal, with the right people assigned. ## Who owns the code, research, and AI workflows? The contract sets ownership, licensing, confidentiality, and handover terms. It distinguishes your product deliverables, research outputs, reusable components, internal team tools, and any automation platform built for you. ## How much does it cost, and when can we start? We quote after reviewing the work, expertise, team size, and availability. Your proposal confirms the price, possible start date, minimum term, and what is included, such as tools, R&D time, support, and handover. ## How do we start a conversation? Email contact@vibehaus.team, or pick a time for an intro call on the homepage. Enter your name and email, then choose any preferred date and time at least 48 hours ahead. The form opens an email draft; we confirm the time by email. ## Related reading - [Fully managed engineering](https://vibehaus.team/managed-engineering) - [Compare delivery models](https://vibehaus.team/compare) - [How we work](https://vibehaus.team/how-we-work) --- # The agentic code factory: from request to releasable change By Vibe Haus · Published and source-reviewed 2026-10-10 Source: https://vibehaus.team/blog/agentic-code-factory A practical design for an agentic code factory: task contracts, isolated execution, evidence, review gates, bounded retries, and release ownership. Original Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed. > A useful code factory produces changes a team can trust. Its unit of work is a releasable change with evidence, an owner, and a recovery path. ## What the factory actually produces An agent can turn a request into a patch quickly enough to make coding look like the whole job. The surrounding decisions remain: what should change, which behavior must survive, whether the patch works in its real environment, and who accepts the consequences of shipping it. Automating the patch without those decisions moves unfinished work into someone else’s queue. Our proposed code factory is a delivery system with explicit states. A request enters as an unresolved problem. It leaves as an accepted change, a rejected experiment, or a documented decision to stop. Each transition requires an artifact another person or tool can inspect. “The agent says it is done” is not a transition condition. Anthropic’s 2024 engineering guidance distinguishes predefined workflows from agents that choose their own next steps. That distinction is useful here: let an agent explore an implementation while ordinary code enforces budgets, permissions, and release gates. The source now notes that its tooling landscape has evolved; this article uses the architectural distinction, not its older tool catalog. ### Workflow: A proposed delivery path. A failed gate never silently becomes permission to ship. 1. Frame: Agree on intent and acceptance evidence. 2. Build: Work within an isolated, bounded task. 3. Verify: Run checks and assemble revision-bound evidence. Decision: Does this revision meet the contract? - Pass → independent review, then authorized release. - Fail → diagnose and retry within budget; otherwise return to the owner. Sources: [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) ## Make the request testable before making it executable Sean Grove’s The New Code argues for preserving structured intent instead of treating code as the only durable record of a change. Our implementation of that idea is a small task contract: the user outcome, constraints, acceptance examples, excluded work, and release owner. It should be short enough to review alongside the diff and specific enough to reject a plausible but wrong implementation. Consider adding CSV export to a customer dashboard. “Add export” leaves permissions, date filtering, column order, large datasets, and spreadsheet formula handling undefined. A useful contract names the authorized user, the records that may appear, the expected output, and the behavior when the export is too large. It also records whether synchronous export is acceptable or a background job is required. Do not demand a perfect specification for uncertain work. Mark unknowns explicitly. A feasibility task can finish with a measured limitation and a recommendation instead of production code. The contract must distinguish an experiment from a commitment to deliver a feature; otherwise a prototype tends to inherit production expectations without production checks. ```yaml # Illustrative task contract; adapt to the actual product. outcome: An account admin can export the current filtered view. acceptance: - Export includes only records the caller may read. - Column order and timezone match the agreed sample. - A second account cannot export the first account's records. unknowns: - Maximum supported export size; measure before choosing sync or async. excluded: - Scheduled reports and third-party delivery. evidence: - Permission regression test and representative output sample. release_owner: Accountable engineer assigned before release. stop_when: - Required data access or product decisions exceed task authority. ``` Sources: [The New Code](https://ai.engineer/talks/8rABwKRsec4-specifications-are-the-new-code) ## Give each worker a bounded place to work Use a separate checkout or worktree for each concurrent code task. Record the base commit and dependencies at creation. A worktree prevents accidental file collisions; it does not isolate processes, credentials, or the network. If untrusted code must run, pair the checkout with an execution boundary appropriate to the risk, such as a constrained container or sandbox. The worker needs the smallest useful set of capabilities: read the relevant repository, edit the intended area, and run the approved checks. A documentation task usually has no reason to receive production credentials. A migration task may need representative schema information without access to customer rows. Keep authority in the runtime and account configuration rather than relying on prose to enforce it. Parallel work is justified when tasks have independent acceptance conditions. Two workers editing the same central schema can create an integration bottleneck even when each patch is locally correct. Split along stable interfaces, publish those interfaces first, or serialize the coupled changes. Account for integration work in the plan instead of treating it as a free final step. - Record base revision, permitted tools, task owner, and maximum execution budget. - Give each task a visible state: ready, running, awaiting decision, verifying, reviewing, or accepted. - Stop and return a concrete question when missing authority or a product decision prevents progress. ## Treat evidence as a build artifact The evidence packet should travel with the candidate change. Include the base and head commits, a concise description of the behavior change, checks actually executed, their results, and important checks that were not run. A browser screenshot can support a layout claim. It cannot establish authorization correctness, database safety, or keyboard behavior by itself. For the export example, useful evidence includes a permitted export, a cross-account denial, a sample output checked against the contract, and a measurement near the agreed size boundary. If the agent only ran unit tests around a helper function, say so. Do not translate “helper test passed” into “export flow verified.” Bind the packet to the exact candidate revision. A later fix, merge, dependency change, or generated-file update can invalidate part of the evidence. Define which checks must rerun when those inputs change. Store durable logs or reports where the reviewer can inspect them; a prose claim without its underlying result is a weaker artifact. - Behavior: what changed and what remained an explicit constraint. - Provenance: base revision, candidate revision, toolchain, and relevant environment. - Verification: commands, results, representative examples, and omissions. - Decision: unresolved risks, assigned owner, and proposed release or stop condition. ## Control work in progress before adding more agents A factory has queues. More implementation workers can increase the arrival rate at review while the review team’s capacity stays fixed. The result is older branches, repeated rebases, and more context switching. Track elapsed time from request to acceptance alongside active work time; a faster implementation stage can coexist with slower delivery. Set a work-in-progress limit at the narrowest stage. For example, a team might initially allow only two review-ready changes per reviewer. That number is an illustrative starting policy, not a universal capacity rule. Adjust it using the observed age of the review queue and the complexity of the changes, rather than maximizing the count of simultaneously running agents. Bound repair loops as well. A failing check should produce a diagnosis before another attempt. Repeated changes that alternate between two failures indicate missing understanding, an unstable test, or a bad task boundary. Stop after the agreed attempt or cost budget, preserve the evidence, and ask the owner to change the task or investigate. A silent infinite loop is an operational failure even if it eventually produces a patch. ## Review and release are separate transitions An independent reviewer should reconstruct the important behavior from the code, requirements, and evidence. Independence means a distinct verification path; merely using a second model does not guarantee it. The reviewer can share the specification with the author while choosing different boundary cases and inspecting the underlying results directly. Acceptance does not automatically grant deployment authority. The release step should use the project’s existing authorization rules, target the reviewed revision, and name the person or automation responsible for monitoring. If the integration branch moved, reconcile that change before reusing earlier approvals. Database or external API changes may need compatibility sequencing rather than a single all-at-once deployment. Recovery must be concrete. A frontend rollback may be simple; a data migration may require a forward repair. Record the signal that would trigger intervention, the owner who watches it, and the viable recovery action. For the export feature, elevated authorization failures or incorrect row exposure matter more than whether the deployment process returned a success code. ## Start with one repeatable task type Choose a task family with a clear outcome and observable checks: a small UI change, a well-scoped bug, or an internal automation. Run the proposed process with one worker and a human reviewer before adding orchestration. Capture where the task stalls and which artifacts the reviewer actually uses. Remove ceremonial paperwork; strengthen missing evidence. Expand only after the team can explain a failed run. Was the request ambiguous, context stale, implementation incorrect, verification incomplete, or release handling weak? These require different repairs. Buying a stronger model may help some implementation failures while leaving all the other failure classes untouched. The process needs five things: contracts, isolated execution, revision-bound evidence, bounded queues, and accountable decisions. Our downloadable task-contract skill is a starting template for the first stage. The companion code-review and evaluation articles explain how to challenge the output and decide whether the overall system is worth expanding. ## Sources and further reading Sources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations. - [The New Code](https://ai.engineer/talks/8rABwKRsec4-specifications-are-the-new-code): Sean Grove · AI Engineer Worlds Fair 2025. Conference talk. - [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents): Anthropic · 2024, with subsequent tooling notice. Engineering guidance. ## Related services - [ai native engineering](https://vibehaus.team/ai-native-engineering) - [managed engineering](https://vibehaus.team/managed-engineering) - [how we work](https://vibehaus.team/how-we-work) ## Reuse You may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms. Downloads and tools: https://vibehaus.team/data --- # Code review after agents: verify the change, not the explanation By Vibe Haus · Published and source-reviewed 2026-10-10 Source: https://vibehaus.team/blog/code-review-for-agents An evidence-first review workflow for agent-written code, covering risk routing, independent tests, actionable findings, and exact-revision release gates. Original Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed. > Agent-written code deserves the same engineering standard as any other change. Faster generation makes review capacity, evidence quality, and risk selection more important. ## Review three things: intent, implementation, evidence A convincing pull request description can make a change feel understood before the reviewer has read the code. Start by separating three things. Intent says what the user needs and what must remain true. Implementation is the actual diff and its surrounding behavior. Evidence shows what was checked against that intent. Agreement between all three is stronger than a polished explanation of any one. Google’s engineering review guidance covers design, functionality, complexity, tests, naming, comments, style, and documentation. Agent involvement does not remove those concerns. Our proposed workflow changes the order: establish the risk and the behavioral claim first, then spend review attention where a plausible patch could still be wrong. For example, an agent may add a permission check to a new export route and describe the route as secure. The reviewer still needs to determine whose account the query uses, whether the caller can choose another account identifier, and whether the underlying data access enforces the same boundary. A correct-looking guard is not evidence that the entire path respects it. ### Workflow: Review attention follows the possible consequence of the change. 1. Route risk: Identify data, permissions, money, availability, and user-visible behavior. 2. Inspect: Read the diff with its callers, requirements, and evidence. 3. Challenge: Exercise a case that could falsify the main claim. Decision: Are material findings resolved on the current revision? - Yes → record the decision and required release checks. - No → return a reproducible finding; review the resulting change. Sources: [What to look for in a code review](https://google.github.io/eng-practices/review/reviewer/looking-for.html) ## Route by consequence, not by author or line count A ten-line authorization change can carry more risk than a thousand-line generated fixture. Classify the consequence of being wrong, the reach of the affected code, and how easily the behavior can be observed or reversed. Record why the route was chosen. Labels such as “small” or “AI-generated” are weak substitutes for this analysis. A low-risk copy change may need a visual check, accessibility sanity check, and content owner review. A payment calculation needs boundary examples, domain expertise, and verification of rounding and retry behavior. A migration needs compatibility and recovery analysis. The policy should explain which evidence is mandatory for each category without forcing every patch through the most expensive path. Automation can propose the category, but ambiguous boundary crossings need a named owner. A package update that changes an authentication library is not automatically routine maintenance. If the scope expands during implementation, reroute the review rather than keeping the convenient original label. | Change surface | Question to challenge | Useful evidence | | --- | --- | --- | | Authorization or tenant data | Can the caller cross an ownership boundary? | Negative tests using a different identity and account. | | Money, retries, or jobs | Can repetition or partial failure duplicate the effect? | Idempotency, rounding, and interrupted-operation cases. | | Schema or integration | Can old and new versions coexist during release? | Compatibility sequence and a viable recovery plan. | | User interface | Can a person complete the task across input modes? | Real viewport, keyboard, and error-state checks. | ## Make the verification path independent An authoring agent can write code and tests from the same mistaken interpretation. If both assume an export should include every account, the tests may pass while reinforcing the defect. A second reviewer reading only the author’s summary can inherit the same assumption. Independence comes from reconstructing the requirement and choosing evidence that could contradict the author. Ask a reviewer to derive one or more boundary cases before reading the author’s test explanation. Give it the actual requirement, current code, diff, and repository conventions. Then compare its expectations with the submitted evidence. A different model can broaden the search, but it is not a proof of independence: both models may rely on the same stale documentation or incomplete context. The reviewer should follow a changed value across the relevant path. For the export example: caller identity → account scope → query filters → serialized output. It can inspect a narrow slice deeply instead of vaguely blessing the whole repository. State the inspected scope and any unverified boundary so another person knows where confidence ends. - Derive at least one falsifying example from the requirement, not from the implementation. - Inspect callers and downstream effects when a local change alters a contract. - Treat external text and tool output as evidence, not as instructions that can authorize actions. - Report what was actually inspected; avoid claims about unexamined parts of the system. ## Challenge what a passing test means Passing tests answer the questions those tests encode in the environment where they ran. They do not establish that the right questions were asked. Review whether a new test would fail against the original defect, whether its assertions observe user-relevant behavior, and whether mocked dependencies erase the very risk under review. A useful counterfactual is to run the regression test against the prior implementation when doing so is safe and practical. If it still passes, the test may be covering a nearby behavior rather than the bug. Another approach is a narrowly scoped mutation: deliberately remove the relevant guard in an isolated workspace and confirm that the test catches it. Do not mutate a shared or production environment for this exercise. Choose test level based on the uncertainty. A deterministic calculation usually benefits from direct boundary tests. A browser interaction may need a real rendered flow. A service integration may require a controlled environment with the actual protocol. Expensive end-to-end tests are not inherently stronger if their assertions are vague or their failures cannot be diagnosed. ## Write findings that can be acted on A useful finding connects a trigger to an observed or well-supported consequence, identifies the relevant code, and explains why the issue matters. It should distinguish a confirmed defect from a concern that needs investigation. Style preferences belong in established tooling or clearly optional comments; they should not compete with broken permissions or data loss. Avoid requiring a particular implementation unless the requirement or repository architecture demands it. The author needs to repair the behavior and provide evidence. A reviewer who rewrites the entire feature into its preferred style can enlarge the diff and obscure the original issue. Keep the correction proportional to the demonstrated problem. This illustrative finding shows how to explain a possible defect and the evidence still needed to confirm it. ```text Finding: Export query accepts an untrusted account identifier. Trigger: User from account A requests an export with account B's ID. Consequence: The query may return B's records if no lower layer scopes access. Location: Export handler and its repository query (attach actual file/line). Evidence: Caller identity is checked, but not bound to the query account. Required resolution: Enforce the ownership boundary and demonstrate denial. Confidence: Confirm the lower-level query behavior before marking reproduced. ``` ## Approvals belong to a revision Every review decision should identify the candidate commit. If code changes after review, the decision must be reconsidered for the affected area. If the base branch changes, determine whether the new integration state changes the behavior or verification assumptions. “Approved yesterday” is not enough information for a release system. Define invalidation rules that are practical rather than ceremonial. A typo correction in a comment may not need a new database compatibility exercise. A lockfile change can alter runtime behavior without changing application source. The release gate should know which checks depend on code, dependencies, generated artifacts, configuration, and environment. The reviewer’s final note should separate findings from release readiness. It can say that no material issues were found within a stated scope while still noting that a required integration environment was unavailable. Missing evidence should remain visible until an accountable person accepts the limitation or completes the check. ## Measure review as part of delivery Track how long review-ready changes wait, how many are returned for substantive corrections, and which defects escape. Do not rank reviewers by comment count or agents by approval rate. Those incentives favor noise or easy tasks. Sample accepted changes periodically to find blind spots, and use incidents to add targeted examples to the review process. When review becomes the bottleneck, reduce the arrival rate or improve task boundaries before adding more authoring agents. Changes that each address one clear behavior can be easier to verify; arbitrarily splitting one coupled feature into many pull requests can make it harder. The useful unit is a change a reviewer can understand and accept with a bounded amount of context. Use the downloadable code-review skill as a starting procedure. It asks for the actual diff, independent checks, revision identity, and actionable findings. It cannot certify security or replace the accountable engineer. Its value is in making a repeatable review standard explicit enough to inspect and improve. ## Sources and further reading Sources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations. - [What to look for in a code review](https://google.github.io/eng-practices/review/reviewer/looking-for.html): Google Engineering Practices. Engineering guidance. ## Related services - [ai native engineering](https://vibehaus.team/ai-native-engineering) - [managed engineering](https://vibehaus.team/managed-engineering) ## Reuse You may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms. Downloads and tools: https://vibehaus.team/data --- # Context engineering: repository knowledge, MCP, and skills By Vibe Haus · Published and source-reviewed 2026-10-10 Source: https://vibehaus.team/blog/context-mcp-and-skills How repository knowledge, task context, MCP tools, and agent skills fit together, with provenance, freshness, trust boundaries, and a practical context contract. Original Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed. > Agents need relevant, current information for the task they are doing. This guide explains how to select that information, provide tools and procedures, and maintain the underlying sources. ## Separate knowledge, procedures, and access An engineering agent needs several different things that are easy to collapse into one large prompt. It needs durable knowledge about the product and repository, temporary facts about the current task, procedures for recurring work, and access to tools or external systems. These layers change at different rates and carry different authority. A repository guide can describe where code lives and which checks are expected. A task contract can identify today’s desired behavior. A skill can explain how to perform a recurring review. The Model Context Protocol (MCP) gives agents a standard way to call tools, such as search or document retrieval. None of those, by itself, establishes that an external document is accurate or that a user authorized a consequential action. The Agent Skills specification defines a discoverable SKILL.md with a name and description, with additional material loaded when useful. MCP defines a protocol for exposing capabilities and exchanging messages. These are complementary mechanisms: a procedure can use a tool, while the tool should retain its own access controls and input validation. ### Workflow: A context assembly path with a freshness and authority check. 1. Locate: Start with the repository map and task contract. 2. Retrieve: Read the smallest relevant source or tool result. 3. Qualify: Record provenance, revision, and unresolved conflicts. Decision: Is the context sufficient and trustworthy for this decision? - Yes → act within the existing authority and preserve evidence. - No → retrieve the missing fact or return the unresolved decision. Sources: [Agent Skills specification](https://agentskills.io/specification); [Streamable HTTP transport specification](https://modelcontextprotocol.io/specification/2025-11-25/basic/transports) ## Keep repository instructions concise and current A repository map should answer where to look next. Name the entry points, ownership boundaries, local commands, and important architectural decisions. Link to detailed documents instead of copying them into a single instruction file. The goal is to help an agent locate the relevant source of truth and recognize when the task crosses into another area. Dex Horthy’s AI Engineer talk presents a research, planning, and implementation approach for complex codebases. We take a narrow lesson from that framing: understanding the current system is a separate artifact worth checking before a broad edit. A plausible plan built on the wrong version of the code remains a bad plan. For a billing task, the map might lead to the account model, invoice calculation, external payment adapter, and migration policy. It should also distinguish maintained decisions from old exploration notes. A workshop proposal and an accepted architecture decision can both be useful, but treating them as equally authoritative produces avoidable contradictions. - Entry points: the few paths that orient a reader to the feature. - Ownership: who resolves product, data, and release decisions. - Verification: commands and environments that check the relevant behavior. - Status: accepted decisions, current implementation, and explicitly tentative ideas. Sources: [No Vibes Allowed: Solving Hard Problems in Complex Codebases](https://ai.engineer/talks/rmvDxxNubIg-context-engineering-for-complex-codebases) ## Build a context contract for the current task Task context should explain the decision being made, the version of the system under discussion, and the boundaries that matter. A file path without a revision can become misleading after a rename or refactor. A log excerpt without environment and timestamp can describe a different deployment. Preserve enough provenance to reconnect a claim to the thing it describes. The contract can remain compact. Include the outcome, relevant source paths, current revision, constraints, acceptance examples, known unknowns, and the next decision. Separate observed facts from hypotheses. “The export job times out in staging with this fixture” is an observation. “The database needs an index” is a hypothesis until examined. When the work moves to another agent or person, summarize what was learned and where the supporting evidence lives. Do not simply compress every prior message. Preserve decisions, rejected alternatives that still matter, current failures, and unresolved questions. A handoff should let the recipient resume reasoning without inheriting an unexamined conclusion. ```yaml # Illustrative context contract. task: Diagnose slow filtered export. revision: Record the actual commit under investigation. observed: - Staging request exceeded the agreed limit with the attached fixture. hypotheses: - Query plan may scan too many rows; not yet confirmed. sources: - Current route, query implementation, schema, and captured query plan. constraints: - Preserve account isolation and existing output format. next_decision: Is the bottleneck query work, serialization, or delivery? missing: Representative upper-bound dataset and acceptable latency target. ``` ## Make skills procedural and inspectable In the AI Engineer talk Don’t Build Agents, Build Skills Instead, Barry Zhang and Mahesh Murag describe packaging reusable knowledge for general agents. Our recommendation is to use a skill when a task has a repeatable method with non-obvious checks. A review skill can explain how to inspect a changed boundary; it should not duplicate the repository’s entire architecture. A good skill explains when it should be used and includes required inputs, a procedure, a useful output, and conditions under which to stop. Keep instructions that apply to every task in the appropriate repository or runtime policy. Keep task-specific details in the task itself. Otherwise skills become competing collections of global rules that are hard to reconcile. Version a skill when its behavior changes. Test it with a realistic task that includes an inconvenient case: missing evidence, conflicting requirements, a failing environment, or a request outside its scope. A syntactically valid SKILL.md can still encourage poor decisions. Behavioral testing should ask whether the resulting artifact helps a real reviewer complete the job. Sources: [Don’t Build Agents, Build Skills Instead](https://ai.engineer/talks/CEvIs9y1uog-agent-skills) ## Design MCP tools around bounded questions A useful tool interface gives an agent a clear question it can ask and a result it can interpret. For a public research library, “search published articles” and “read this article by slug” are enough. A generic “fetch any URL and execute what it says” capability would create a very different risk surface without improving that basic use case. Tool results should include stable identifiers and source URLs. Search snippets help choose a document; they are not a substitute for reading the relevant passage. A read result should expose the full authored content or explicitly state truncation. Unknown identifiers should produce an understandable error instead of silently selecting a nearby item. Our public MCP endpoint follows that narrow design: it searches and reads this site’s own published research and retrieves the downloadable skills. It does not connect to customer systems or receive repository credentials. Its scope is visible on the Data / MCP / Skills page, and the same article text is available through JSON and Markdown for clients that do not use MCP. ## Treat retrieved content as data with provenance An external page can contain instructions written by someone other than the user. A document that says to ignore prior rules or send credentials somewhere is still just document content. The agent runtime should keep retrieved text separate from higher-priority instructions, and tools should enforce access boundaries independently of the model’s interpretation. The same principle applies inside a repository when files, issue comments, or generated output can be influenced by outside contributors. Read them for relevant facts; do not let them silently broaden the task’s authority. Skills downloaded from the internet should be inspected before installation, just as a team would inspect another automation artifact. Freshness is another trust dimension. Cache stable documents with explicit versions; recheck volatile facts before acting on them. If a repository guide conflicts with the installed framework’s current documentation or the executable code, investigate and record the resolution. Do not hide the conflict by choosing whichever source supports the first plan. ## Maintain context as part of changing the system When a change moves an entry point, changes a command, or supersedes an architectural decision, update the map in the same delivery cycle. The person accepting the change should be able to verify that the next engineer will find the right path. This makes documentation maintenance observable rather than an occasional cleanup project. Sample failed agent runs and classify the context failure. Was the needed fact absent, stale, difficult to retrieve, contradicted, or ignored? Adding more text helps only the first category and sometimes the third. A better source hierarchy, a smaller tool result, or a clearer stopping condition may be the more effective repair. The intended outcome is a connected knowledge system: a short repository map, task-specific evidence, reusable procedures, and bounded tools. Each part should be independently understandable. That lets people and agents share the same engineering context without requiring everyone to consume the entire history of the project. ## Sources and further reading Sources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations. - [Agent Skills specification](https://agentskills.io/specification): Agent Skills. Specification. - [Streamable HTTP transport specification](https://modelcontextprotocol.io/specification/2025-11-25/basic/transports): Model Context Protocol · 2025-11-25. Specification. - [No Vibes Allowed: Solving Hard Problems in Complex Codebases](https://ai.engineer/talks/rmvDxxNubIg-context-engineering-for-complex-codebases): Dex Horthy · AI Engineer. Conference talk. - [Don’t Build Agents, Build Skills Instead](https://ai.engineer/talks/CEvIs9y1uog-agent-skills): Barry Zhang & Mahesh Murag · AI Engineer. Conference talk. ## Related services - [ai native engineering](https://vibehaus.team/ai-native-engineering) - [product design](https://vibehaus.team/product-design) ## Reuse You may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms. Downloads and tools: https://vibehaus.team/data --- # Evaluating agentic engineering: quality, cost, and delivery By Vibe Haus · Published and source-reviewed 2026-10-10 Source: https://vibehaus.team/blog/evaluating-agentic-engineering A measurement plan for agentic engineering that includes accepted outcomes, human review, rework, failures, cost, and the limits of current productivity evidence. Original Vibe Haus synthesis and proposed engineering workflows, informed by the linked primary sources. These articles are not peer-reviewed studies or measured customer results. Examples and thresholds are illustrative unless explicitly attributed. > The question is whether the team delivers more valuable, dependable work for the resources it spends. Token speed, generated lines, and convincing demos cannot answer that alone. ## Read productivity claims in their actual setting METR’s early-2025 randomized study followed 16 experienced open-source developers working on 246 tasks in repositories they knew well. With the tested AI tools available, tasks took 19% longer on average. This is useful evidence about that population, task mix, and tool period. It is not a universal estimate for every developer, greenfield product, or later agent. The February 2026 follow-up matters too. METR reports that selection effects and difficulties measuring concurrent agent work make its newer productivity estimates unreliable. Developers and tasks opting out of the no-AI condition can change who remains in the experiment. The authors suggest improvements are plausible while warning that their data gives weak evidence for the size of the change. DORA’s 2025 report frames AI as amplifying existing organizational strengths and weaknesses. Treat that organizational research as context for adoption, not as a causal guarantee for your team. The practical response to mixed and changing evidence is to measure your own work carefully, while keeping the limits of that measurement visible. Sources: [Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/); [We are Changing our Developer Productivity Experiment Design](https://metr.org/blog/2026-02-24-uplift-update/); [State of AI-assisted Software Development](https://dora.dev/research/2025/dora-report/) ## Define the work being measured A completed coding task is not necessarily a delivered product outcome. Decide whether the experiment concerns implementation speed, accepted changes, released features, or customer results. Each requires a different observation window. A benchmark patch can be evaluated before release; customer retention cannot. For an initial engineering pilot, use a change with a clear outcome that a reviewer can accept. Define completion in advance: required behavior, checks, review disposition, and unresolved limitations. Include failed and abandoned attempts in the denominator. Otherwise a system that spends heavily on ten attempts and succeeds once can appear equivalent to one that succeeds immediately. Segment the task set. A documentation correction, an unfamiliar integration, and a concurrency bug exercise different capabilities. Record repository familiarity, task uncertainty, risk, and dependency complexity. Report performance by meaningful segment before combining it into an average. A gain on easy tasks can conceal a regression on the work that occupies most of the team. ### Workflow: A proposed evaluation loop. Keep a held-out set separate from workflow tuning. 1. Define: Choose tasks, acceptance rules, and the decision to be made. 2. Compare: Run a baseline and candidate under recorded conditions. 3. Inspect: Score quality, total effort, failures, and cost. Decision: Does the candidate meet the predeclared quality and cost criteria? - Yes → expand gradually and keep monitoring production outcomes. - No → diagnose by task segment, change one cause, then reevaluate. ## Build a small, representative task set Start with real task shapes from the team’s backlog or recent history. Historical tasks are useful only if the evaluation prevents the candidate from seeing the finished solution. Freeze the starting repository revision, remove answer-bearing artifacts from the candidate’s context, and check whether the model or tool could retrieve the solution elsewhere. Report contamination you cannot exclude. Write acceptance criteria independently of the candidate implementation. Include ordinary paths, relevant boundary cases, and at least some tasks that require recognizing insufficient information. An agent that stops for a missing authorization decision may be behaving better than one that produces a complete-looking patch by inventing the answer. Keep a development set for tuning prompts, tools, and skills, and a separate held-out set for decisions. Repeatedly optimizing against the same examples can produce a workflow that is good at the evaluation rather than the job. Version the tasks and record when a task becomes unsuitable because the product or tooling changed. - Preserve a reproducible starting state and the available context. - Specify acceptance evidence before seeing the candidate result. - Include failures, ambiguity, and realistic constraints. - Keep held-out tasks out of routine prompt and skill tuning. ## Compare workflows and document the limits A baseline can be the team’s current workflow, a simpler agent setup, or a different model configuration. Choose the comparison that answers the adoption decision. If the question is whether an orchestration layer helps, hold the model and task conditions as stable as practical instead of changing everything at once. For task-level comparisons, randomize assignment where feasible and balance important task categories. Repeating the same task with the same developer introduces learning; paired tasks need to be comparable rather than identical in a way that leaks the solution. Track deviations from the assigned workflow and report exclusions explicitly. Small pilots provide directional evidence, not a precise universal effect. Agents introduce variability. Record model identifier, tool versions, reasoning settings when available, skills, repository revision, resource limits, and number of attempts. A single successful run cannot establish reliability. Repeat enough cases to see whether the result depends on a lucky attempt, and report the distribution rather than only the best outcome. ## Count the work that generation metrics omit Capture human time spent framing, supervising, reviewing, integrating, and repairing. Capture model, tool, and infrastructure cost separately. Record wall-clock lead time as well: a task can consume little active human time while waiting several days in a review queue. Parallel agents make these measures diverge even more. Use accepted outcomes as the denominator for an economic comparison, while keeping quality thresholds fixed. The illustrative example below assumes the same task mix and acceptance standard. It is arithmetic to show the method, not a measured Vibe Haus result or a promise about an agent system. Here the candidate uses less total money but produces fewer accepted changes. Cost per accepted change therefore rises from $100 to about $105.56. The result could still be acceptable for another reason, such as shorter lead time, but that reason must be measured rather than inferred from the smaller total bill. ```text cost per accepted change = (human effort cost + model cost + tools + infrastructure) / accepted changes Report separately: - unresolved rework and the observation window - quality failures and severity - request-to-acceptance lead time - tasks excluded or abandoned ``` | Illustrative measure | Baseline | Candidate | | --- | --- | --- | | Attempted changes | 12 | 12 | | Accepted changes | 10 | 9 | | Human effort at $100/hour | 9 hours = $900 | 7 hours = $700 | | Model, tools, infrastructure | $100 | $250 | | Total measured cost | $1,000 | $950 | | Cost per accepted change | $100 | $105.56 | ## Define acceptance thresholds before the trial Predeclare the conditions that would stop or limit rollout. These might include an unacceptable permission defect, a cost ceiling, excessive reviewer intervention, or unreliable completion on a critical task family. The thresholds belong to the product’s risk and economics. Do not choose them after seeing the results to make a preferred tool win. Combine deterministic checks with independent judgment where necessary. Tests can establish explicit behavior; expert review can examine architecture, maintainability, and missing cases. Reviewers should use a shared rubric and record disagreements. When practical, hide which workflow produced a candidate to reduce expectation effects, while recognizing that style or artifacts may reveal it. Track escaped defects over an appropriate window after acceptance. A pilot that ends at merge cannot measure downstream maintenance or incidents. Keep those outcomes separate from immediate pass rate, and describe the lag. The absence of an observed incident in a small short trial is weak evidence about rare high-impact failures. ## Turn the result into an operating decision An evaluation should end with a specific decision: adopt for a named task family, continue a bounded pilot, revise the workflow, or stop. Explain what evidence would change that decision. A partial success can justify a narrow deployment while leaving higher-risk work under a different process. If the candidate fails, diagnose the stage. Ambiguous tasks call for better contracts. Repeated tool misuse calls for better interfaces or constraints. High review effort may indicate broad diffs or weak evidence. Expensive retries may point to an unsuitable task class. Changing the model is one possible intervention, not the default explanation for every failure. The downloadable evaluation skill turns this article into a reusable plan and report structure. Use it before a pilot to make the comparison fair, then again afterward to identify missing data and limitations. The aim is a decision the team can defend with observed outcomes, including the inconvenient ones. ## Sources and further reading Sources reviewed 2026-10-10. This is a focused reading list, not an exhaustive literature review. Source dates and study conditions matter; follow the original links for their full methods and limitations. - [Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/): METR · July 10, 2025. Randomized study. - [We are Changing our Developer Productivity Experiment Design](https://metr.org/blog/2026-02-24-uplift-update/): METR · February 24, 2026. Study update. - [State of AI-assisted Software Development](https://dora.dev/research/2025/dora-report/): DORA · 2025. Research report. ## Related services - [research and development](https://vibehaus.team/research-and-development) - [engineering cost](https://vibehaus.team/engineering-cost) ## Reuse You may use and adapt our original templates and workflows with attribution. Third-party sources retain their own terms. Downloads and tools: https://vibehaus.team/data