A test system that can say I don't know is valuable
Explore Derek Wang's AI harness engineering insights on why error ledgers and knowing unknowns matter more than simple pass results.

Stock photo for illustration only, not from the actual event
- Karl Popper argued science advances by eliminating errors rather than accumulating truths.
- AI test systems need failure ledgers instead of simple pass/fail gates.
- Full regression runs take 10 seconds and scale across 50 scenarios in 6 suites.
- Real-world usage successfully intercepted 22 hidden HTTP-500 errors before shipping.
Karl Popper spent the middle of the twentieth century dismantling the traditional view of scientific accumulation. He argued that science does not advance by proving things right more often, but by tearing wrong answers out of the game in broad daylight. A beautiful theory is worth only what it survives, and every knocked-down guess removes a false possibility from the board through guess-and-refute testing.
Harnessing an AI system requires building a similarly rigorous and disagreeable machine for code. Rather than creating a machine that only proves code is correct, developers must build a system that relentlessly tries to prove it wrong while maintaining a ledger of every historical failure. This failure ledger addresses a critical void where most growing AI architectures currently have nothing at all.
Most software teams begin testing in an identical, unstructured manner where an agent checks an output and declares it valid without any underlying baseline or accountability. Without counting past failures, a team has no accurate measure of its accumulated technical debt.

Stock photo for illustration only, not from the actual event
The full regression suite eventually ran in just ten seconds, a timeframe fast enough to be skipped—which is precisely how confidence quietly turns into a falsehood. Speed and safety operate as the exact same lever here; rapid feedback enables bolder correct changes while discouraging incorrect ones.
Achieving a 10-second test execution cycle directly aligns with the fast feedback loop principle in modern software engineering. When autonomous agents receive rapid validation, they can safely refactor complex architecture layers without fearing undetected regressions, significantly boosting iteration velocity and overall system resilience.
Framing verification as a ledger rather than a gate changes team dynamics completely. While a gate merely halts bad work, a bridge allows productive work to travel safely, empowering AI agents to refactor code with immediate awareness of potential systemic impact rather than shooting blindly into the dark.
Treating failure data as a growing organizational asset shifts testing from compliance checking to cartography of system vulnerabilities. This historical context prevents repetitive debugging cycles across project iterations.
"A test system that can say 'I don't know' is worth more than one that says 'passed'"
Derek Wang
The fulltest framework integrates three essential mechanisms absent from standard runners: classifying failures before judgment using historical ledgers, growing the ledger organically from real incidents and git commits, and raising the operational baseline automatically so the system cannot silently backslide.
Ultimately, a massive test suite means very little if it measures the wrong metrics or never interacts with actual failure paths. The only metric that truly matters is how many past real incidents the suite would have successfully caught. When applied in practice, a 10-second run scaled into 50 scenarios across six specialized suites, successfully intercepting 22 hidden HTTP-500 errors before release.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment