← all articlesDeutsche Fassung
Vibe coding vs. spec-driven development: same result, better code
The same application built three times with Claude Code, once in conversation and twice from a specification. For this application of medium complexity, vibe coding was no worse in outcome. But the code of the spec version is split along domain boundaries, and that matters for maintenance and further development.
I had Claude Code build the same application three times: once in conversation and twice from a specification that the agent drafted from my briefing and I approved. The result was practically the same, the code was not. For this application of medium complexity, vibe coding got me just as far. But only the spec versions are split along domain boundaries, and that decides how well an application can be maintained and developed further. I also noticed that the agent asked questions while writing the specification. When vibe coding, it did not.
Vibe coding means something different today
Andrej Karpathy coined the term in February 2025: accept all of the AI's suggestions, stop reading the changes, in his words good enough for throwaway weekend projects. Nine months later it was Collins Word of the Year. An agent like Claude Code with Opus 5.5 now writes its own tests and sticks to a project's conventions.
Still, my hypothesis was that vibe coding stops working beyond a certain complexity. As an application grows, it gets harder for people and agents alike to keep track of the whole. Terms mean different things in different places: a "visit" is something different for reception than for the person who sends the invitation. And a change in one place breaks something in another. That is what domain-driven design is for. You split the domain, the subject area in question, into subdomains. With bounded contexts you draw boundaries within which a model and its terms have one clear meaning. That way each part can be understood and changed on its own. My assumption was that you have to give an agent this decomposition and describe the data model formally. Tools for spec-driven development such as GitHub's Spec Kit or Amazon's Kiro go in this direction: first the specification, then the code. For a few weeks I have been working along the lines of Anthropic's playbook for an AI-native development lifecycle.
Then came a post by Karpathy on X on October 2, 2026. He recommended having language models answer in ASD-STE100, STE for short, Simplified Technical English, a controlled language from aircraft maintenance. I was interested in the opposite direction: does it help if the requirements for the agent are written that way, or in EARS, the sentence template Kiro uses?
The experiment
The application is a visitor registration system, for which we had actually planned to buy off-the-shelf software, built as a module of an internal platform: employees register visitors and groups, and reception sees who is expected. Medium complexity, not a business-critical system.
The vibe version started with my briefing, about one page written from memory, after which I steered in conversation. For the other two, the agent drafted a specification from the same briefing, once with requirements in EARS and once in STE. I answered its open questions, approved the specification and plan, and tried out the result myself along the way. I measured with 24 acceptance scenarios that an agent had written from my briefing and my preliminary notes before the first build. I only read them once the vibe version was finished. 17 follow from the briefing, 7 only from preliminary notes that none of the versions got to see, such as an emergency list for reception. An agent played them through in the browser, and I spot-checked the failed ones. After that, Vibe and EARS each got the test report once to fix things, and finally a larger change request with 13 new scenarios. I measured STE only at the first acceptance.
The results
| first acceptance | after the test report | change request | |
|---|---|---|---|
| Vibe | 20 of 24 | 24 of 24 | 13 of 13, the old 24 still passed |
| Spec with EARS | 19 of 24 | 24 of 24 | 13 of 13, the old 24 still passed |
| Spec with STE | 19 of 24 | not measured | not measured |

Vibe owes its one-point lead to an edge case that the testing agent marked as failed and I counted as passed. In practice it is a tie. So vibe coding worked for this application. The agent wrote around 300 tests for the module and followed the platform's conventions.
What surprised me most was the pace. Nine of my prompts adjusted the product, for example the retention period or that the security desk starts directly at reception. In the spec version, too, a trial run brought up two such points. In the past I often only saw things like that in the next sprint review. Here there were about eleven minutes between two prompts on average, including my time to read and think. Much of what is in no briefing only occurs to you once you have the application in front of you.
The code still shows a difference. The vibe version has around 6,300 lines and technically follows the platform's template. In terms of the domain it is not split up: most of the domain logic for eight functional areas sits in one file with 1,687 lines. The next largest file of this kind in the platform has 350. The EARS version gets by with around 4,300 lines, with slightly fewer features. Its specification splits the domain into subdomains: visit registration and reception as core, events as supporting, access and staff directory as generic, each with its own bounded context. The new contexts show up in the code as separate packages, even if their boundaries are bypassed in two places. In the vibe version the data model exists only as a database schema and as the input formats of the user interface; the EARS version describes entities, attributes and rules explicitly.

The EARS version is not more bug-free. Two agents reviewed the versions blind and found bugs in both that no scenario had uncovered.
Why the split matters for maintenance
That loosely coupled parts with high internal cohesion are easier to change is one of the oldest principles of software architecture. David Parnas described the idea behind it in 1972: a module should hide a decision so that it can be changed without touching the others. Stevens, Myers and Constantine made these properties known as coupling and cohesion in 1974.
The split matters for a coding agent, too. Whatever it has to read for a change takes up space in its context window, the model's limited working memory. If something about events needs to change in the EARS version, it can in principle make do with one package. In the vibe version it has to find its way around a file with all eight functional areas, and every extension makes that file bigger.
The one change I measured did not show this yet. Both versions needed 3 prompts, and the spec version even used more tokens because it created a whole cycle of documents for it. At this size the agent can probably still keep track of the large file. For me the split is still the most important property of the spec version: an application grows with every change, and the worse it is split, the more expensive the next one gets.
EARS or STE made no difference
EARS and STE passed exactly the same scenarios at the first acceptance, as they already had in a small pilot with a simpler task. I did not compare a specification written in free language. So the experiment does not show whether the form of the requirements matters at all.
Spec-driven development brings back an old hope: that a formal model is the real work and the code follows from it. Model Driven Architecture promised that in the 2000s, and now some recommend ontologies as a foundation for agents. In my experience little of it caught on, because the model next to the code became a second truth that nobody maintained. Today the agent can maintain it. Who reads it is another question. I, too, only skimmed the specification. In my experiment, what helped was a good split along the domain and the questions the briefing left open. A described data model is part of that; a model that code is generated from was not needed.
The agent asks when it writes a specification
Of the 7 points that were only in my preliminary notes, Vibe had 3 and the spec versions 2 each. None of them had the emergency list. Even Opus 5.5 can't read minds.
While writing the specification, on the other hand, the agent named 16 open questions and one ambiguous passage in my briefing and suggested an answer for almost every one. Some of it I had not thought of. For example, it noticed that my briefing leaves open whether the security desk should see the contact details of the visitor or of the person being visited, and it asked how long visitor data should be kept. The vibe version also added things I had not asked for. The agent wrote them into a description of the module, but it did not ask beforehand.
I accepted most of the suggestions. I decided differently on whether group names must be unique and, only while trying things out, on the retention period. One of the questions was whether registrations should be editable afterwards. The suggestion was "not in this cycle", and I agreed. Today I miss it. In none of the versions can I add visitors to an existing registration.
I only noticed this late. An agent had written the change request in advance for the experiment, and I did not know it until the measurement. So nobody checked it from a business perspective. It introduced seat limits with a waiting list, which we don't have in that form. For the comparison it remains fair, but it missed the actual need.
What the numbers don't say
This is a case study with one run each. The agent drafted the specification, and in terms of the domain it rests only on my briefing and my answers. I only skimmed the specification, plan and scenarios for domain correctness. So my hypothesis about a decomposition by people with domain knowledge has not been tested. The spec versions were built after the vibe version, when I already knew the domain and the acceptance scenarios, and everything ran within one week. And with the platform's conventions, the vibe version itself had a kind of specification, not for the domain, but for the technology. I am not publishing code and data because the application is used internally.
Conclusion
In my experiment the agent largely took over planning and testing itself. Two things remain important: how the application is split, and the answers to the domain questions. The specification helped with both in my experiment. It split the domain into bounded contexts and made open questions about my briefing visible. Whether the requirements were in EARS or STE made no difference, and a model that code is generated from was not needed.
Sources
- Andrej Karpathy on X, February 2, 2025 (vibe coding)
- Collins: Word of the Year 2025
- Andrej Karpathy on X, October 2, 2026 (ASD-STE100)
- GitHub: Spec Kit
- Kiro: Feature Specs
- Anthropic: The AI-native SDLC playbook
- Easy Approach to Requirements Syntax (EARS)
- Simplified Technical English (ASD-STE100)
- David L. Parnas: On the Criteria to Be Used in Decomposing Systems into Modules, 1972
- W. P. Stevens, G. J. Myers, L. L. Constantine: Structured Design, 1974
- OMG: MDA Guide, Version 1.0, 2003
- CIO: The next enterprise architecture asset: ontologies for AI, May 2026