Executive summary
A fast prototype can leave verification, integration, security, recovery, and maintenance unresolved. Where does the business cost move?
The demo changes the budget conversation
A working screen can change a business conversation overnight. An idea that needed a development estimate now looks like something a team could launch next week. AI assistance makes it easier to explore that idea, test an interface, or assemble an initial workflow. That is useful progress. The difficult question is what the demonstration actually proves.
Imagine a dispatch tool that lets an operator assign a delivery and notify a customer. In this hypothetical example, the demonstration handles one assignment beautifully. It says little about two operators assigning the same job, a notification service timing out, or a customer seeing another customer's details. Those questions exist whether a person or an AI wrote the implementation.
For a business, the cost of software includes the work needed to trust its behavior and keep it useful. A fast initial build can leave that work outstanding. Treating the first demonstration as a launch commitment makes the remaining work feel like an unexpected expense, even when it was always necessary. The useful budget question is therefore: what has become cheaper to create, and what still needs to be proven?
Name the workflow, not the tool
Here, vibe coding means a prompt-led approach in which someone describes a desired outcome, accepts generated changes, and steers through feedback while minimizing their understanding or review of the implementation. That can be a reasonable way to explore an idea within controlled boundaries. It becomes a different proposition when the result starts serving customers or making consequential decisions.
AI-assisted engineering is broader. An engineer can use a model to investigate unfamiliar code, propose tests, compare designs, or implement a bounded change while retaining responsibility for the system. The distinction concerns how decisions are checked and owned. Using Codex, Claude Code, Copilot, Cursor, or another coding assistant does not by itself make a workflow vibe coding.
This article's starting point is Addy Osmani's public Beyond Vibe Coding framing, which distinguishes these approaches. It inspired this Evolium analysis; it does not imply that Osmani endorses Evolium or this series.
A progress estimate can conceal unfinished decisions
Osmani describes a 70% problem: reaching a persuasive prototype can feel much easier than completing the work needed for dependable use. The 70% / 30% framing is a practical heuristic, not a universal measured ratio. It cannot tell a business how much of its particular project is complete, how long delivery will take, or what the remaining work will cost.
Its value is in questioning apparent progress. Screens and successful demonstrations are visible. Decisions about access, failure, and maintenance are less visible until someone asks for evidence. A team may have completed most of what a viewer can see while leaving important operating decisions unresolved. Count those decisions separately from the number of features on screen.
What must be settled before people depend on it?
Start with the business promise. Who may use the system, what may they do, and what outcome counts as correct? Acceptance criteria should describe inconvenient cases as well as the ideal path. For the imagined dispatch tool, an assignment should have a defined owner even if two operators act together. A polished confirmation message is insufficient evidence that this rule holds.
Then establish boundaries. Decide which component owns the assignment, which service sends the notification, and where the authoritative record lives. Define what happens when those components disagree. Architecture earns its keep here by making responsibility and failure behavior understandable, rather than by making a diagram elaborate.
Data handling belongs in the same discussion. Review authentication, authorization, secrets, retention, and the information available to each actor. GitHub's Copilot Chat application card warns that apparently plausible output may still be incorrect and recommends review and testing, including security review. This is responsible-use guidance, not a measurement of how often generated code fails.
The development environment also needs boundaries. OWASP's AI secure-coding guidance explains that coding agents may read files, run commands, install packages, or use networks, depending on their capabilities and permissions. A file excluded by .gitignore may still be readable by an agent; version-control exclusion does not establish a context boundary. Check actual access and context controls before allowing secrets or sensitive information into a workflow. Agent configurations differ, so assess the tool being used rather than assuming identical powers across products.
Integration creates another set of obligations. Check external API behavior, dependency provenance, credential handling, and compatibility. An unavailable service needs a defined response. Retrying a request needs a rule that prevents an unintended duplicate action. A dependency update needs someone who can evaluate its effect. AI can help investigate these questions, but a generated answer still needs verification against the actual system.
Verification should include real operating paths. Test delayed responses, interrupted sessions, empty results, and unauthorized requests. Exercise keyboard use, narrow screens, and the complete browser workflow. Google's web engineering codelab discusses gaps in low-context, single-prompt workflows and the importance of browser verification alongside tests. It offers engineering guidance, not a quantified failure rate.
Finally, make operation and future change possible. Decide what signals reveal a problem, who receives them, and how an operator can recover or reverse a release. Keep enough documentation to explain the system's decisions and constraints. Name the people responsible for maintenance and for approving consequential actions. These are engineering activities AI can assist with; their completion must be demonstrated by the team that owns the result.
The costs show up in different places
The first cost category is rework. If access rules or data relationships are discovered after the interface is polished, the team may need to change several layers together. Launch can then be delayed after an apparently fast prototype. Neither outcome is inevitable, but a plan that omits unresolved requirements has no allowance for resolving them.
Integration fragility can become an operating expense. A workflow that only handles a successful upstream response may require manual intervention when a supplier changes an API or a service becomes unavailable. Interruptions consume attention from the people whose work the software was meant to support. The business needs to understand the fallback before relying on the automation.
Security and data exposure are risk categories, not automatic consequences of using AI. Evaluate the actual permissions, data movement, dependencies, and deployment design. An unexamined package can add a supply-chain obligation; an unclear access rule can create exposure risk. Human-written software also requires these checks. The origin of a line of code does not certify its safety.
Maintenance carries its own burden. Someone must understand enough to diagnose an issue, make a change, and assess what else that change affects. When ownership is unclear, a seemingly small request may need fresh investigation before any useful work begins. A business can inherit this uncertainty even if the initial demonstration was inexpensive.
False confidence ties these categories together. A passing test proves what that test exercised. A demo proves the demonstrated path. Neither alone establishes production readiness. A better progress report identifies what has been verified, what remains uncertain, and which decisions block dependable use. That gives the business a basis for scheduling and funding the remaining work without inventing a financial return.
What the evidence can support
In METR's early-2025 randomized trial, 16 experienced open-source developers completed 246 tasks in mature repositories they knew well. With early-2025 AI tooling allowed, completion took 19% longer in that specific setting. The result concerns those developers, tasks, repositories, and tools; it does not establish that AI slows every developer or every kind of work.
METR's February 2026 update says its newer experiment could not yield a reliable current productivity signal because participation and task-selection effects had become problematic. Changed participation incentives and difficulty measuring time with concurrent agents also complicated interpretation. The 2025 result therefore cannot serve as a universal 2026 productivity estimate.
There is positive evidence too. A GitHub-published randomized study of an API implementation exercise reported better unit-test outcomes and code-review ratings for the Copilot group. This is vendor research in a bounded task, rather than independent proof of a benefit for every project. It provides a useful counterweight to a claim that AI assistance necessarily reduces quality.
These sources examine different questions. A completion-time study, a code-quality exercise, and practical engineering guidance are not interchangeable measures. Our interpretation is that a team should evaluate its own task and process: time to a verified outcome, effort spent reviewing and correcting it, and the ability to maintain the result. Context, verification, and practitioner judgment matter more than a universal claim that AI is faster or slower.
Set the boundary by consequence
For exploration, a limited understanding of the implementation can be acceptable when the experiment is disposable and tightly contained. Interface sketches, prototypes using synthetic data, and isolated proof-of-concept workflows can help a team learn what it wants. Keep their credentials, connections, and promises limited enough that an unexpected result does not silently affect real operations.
Raise the engineering standard when the software serves customers, manages authentication or authorization, moves money, handles confidential or regulated information, or connects to core operations. Long-lived systems that another team must maintain also need explicit design and ownership. Automated actions with material consequences require a person or team accountable for the decision boundary and its verification.
An internal label does not make a tool low-consequence. An internal experiment with access to customer records or the ability to change orders may warrant substantial controls. This is an engineering decision framework, not legal or compliance advice. Evaluate the effects the system can produce, the data it can reach, and the difficulty of correcting an error.
A release conversation worth having
Before promoting a prototype into dependable use, ask for concrete answers:
- Promise: Can the business and engineering owners agree on acceptance criteria, including cases that should be refused?
- Authority: Is there a documented map of data access, permitted actions, and what the AI assistant may change or execute?
- Responsibility: Who can explain the code and system, approve consequential changes, and maintain them after the original builder leaves?
- Evidence: Have tests and real user workflows, including browser and accessibility checks, demonstrated the agreed behavior?
- Connections: Have dependencies, credentials, authorization rules, retries, and external-service failures been reviewed together?
- Recovery: Can the team detect a problem, identify the affected work, and rehearse a safe rollback or recovery?
- Operating fit: Has someone validated representative working cases and exceptions with the people who will use the system?
- Handoff: Are unresolved risks, change limits, and maintenance obligations recorded, with an owner and a decision for each?
The checklist is a discussion about evidence, not a certificate. Its answers should influence the launch decision. Where an answer is missing, either resolve it or narrow the system's reach until the remaining uncertainty is acceptable to its accountable owner.
Cheaper generation makes judgment more valuable
AI assistance can lower the effort needed to generate software and make useful experiments accessible earlier. The business opportunity is to turn that speed into learning and verified delivery. Skipping the decisions that make software understandable and operable merely leaves those decisions for later.
Research. Engineer. Evaluate. Govern. Evolve. That sequence is useful here because it keeps creation connected to evidence and continuing ownership. A business should optimize for a system it can understand, verify, operate, and change. The strongest outcome is a useful capability whose behavior someone can stand behind after the demonstration ends.
Limitations and uncertainty
This material provides a general decision framework. Usefulness, risk, and implementation depend on each organization’s objective, systems, data, responsibilities, and constraints.
Practical next step
Define the operating problem, identify the decision owner, and document the evidence required to evaluate an alternative.
Beyond the Trend / 01
Continue in the series
Separating durable operational value from technology hype, and examining what emerging technology means for real businesses and systems.
