Fixing Spec Rot and Bloated AI Workflows in SDD
An analysis of spec rot caused by tools like OpenSpec and spec-kit, generating excessive markdown files, and a lean approach to SDD.

Stock photo for illustration only, not from the actual event
- Spec-driven development often suffers from spec rot due to unwanted automated knowledge bases.
- Scott Logic recorded 2,577 lines of generated markdown versus just 689 lines of code.
- Managing context budgets and limiting unnecessary artifacts keeps codebases clean and lightweight.
- Switching to minimal setups like Pi and a 9-line AGENTS.md eliminates software bloat.
When mentioning spec-driven development (SDD) out loud, the inevitable response is often the mention of spec rot. Tools like OpenSpec, spec-kit, and BMAD have earned this reputation by building accumulated knowledge bases that developers never asked for, turning documentation into long-term overhead.
SDD consists of two primary steps: writing the specification and building directly from it. Both steps operate under strict context budget constraints, considering that the 1M window remains largely a marketing metric. Forcing AI models to read through massive legacy logs and generated files session after session wastes significant processing capacity.
A notable test by Scott Logic running spec-kit on a single feature generated 2,577 lines of markdown compared to just 689 lines of actual code, alongside 3.5 hours of review time. Before this implementation, deleting roughly 250MB of accumulated memories, plans, task lists, and session leftovers from ~/.claude caused zero breaks and went completely unnoticed by the system.

Stock photo for illustration only, not from the actual event
This highlights the problem of numerous abstraction layers claiming to remember information without any single true owner. Matters like bracket placement ({) belong to code formatting tools like make fmt, rather than requiring searches across multiple trackers, memory stores, and ADR documents.
"A constitution, a spec, a plan and a task list per feature, a memlog. Scott Logic ran spec-kit on one feature and counted 2,577 lines of generated markdown against 689 lines of code, plus 3.5 hours of review."
Scott Logic
The solution is to stop mixing unrelated responsibilities. Sticking to the classic advice of one tool for one job works exceptionally well here. Transitioning to Pi removes all software bloat and unwanted opinions, maintaining a global AGENTS.md file consisting of just nine lines: six invariants, two headers, and one environment note.
The phenomenon of spec rot highlights how over-relying on automated AI agents and large language models can introduce architectural complexity. While generating massive amounts of automated documentation sounds appealing initially, these intermediate artifacts accumulate into technical debt that burdens both human developers and models. Returning to a minimalist approach where the codebase remains the sole source of truth is far more sustainable.
Operational workflows are strictly divided: research handles read-only queries by splitting questions and returning cited findings without writing files. Transfer handles markdown handoffs with redacted secrets, and Apply verifies claims against the repository to execute minimal changes. Crossing the 100k token threshold triggers an immediate session transfer and restart.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment