Skip to main content

A test system that can say I don't know is valuable

Explore Derek Wang's AI harness engineering insights on why error ledgers and knowing unknowns matter more than simple pass results.

AI-written
Inewgen
28 Sep 2026Source: Dev.to4 min read (0 views)
Share
A test system that can say I don't know is valuable

Stock photo for illustration only, not from the actual event

Font size
  • Karl Popper argued science advances by eliminating errors rather than accumulating truths.
  • AI test systems need failure ledgers instead of simple pass/fail gates.
  • Full regression runs take 10 seconds and scale across 50 scenarios in 6 suites.
  • Real-world usage successfully intercepted 22 hidden HTTP-500 errors before shipping.

Karl Popper spent the middle of the twentieth century dismantling the traditional view of scientific accumulation. He argued that science does not advance by proving things right more often, but by tearing wrong answers out of the game in broad daylight. A beautiful theory is worth only what it survives, and every knocked-down guess removes a false possibility from the board through guess-and-refute testing.

Harnessing an AI system requires building a similarly rigorous and disagreeable machine for code. Rather than creating a machine that only proves code is correct, developers must build a system that relentlessly tries to prove it wrong while maintaining a ledger of every historical failure. This failure ledger addresses a critical void where most growing AI architectures currently have nothing at all.

Most software teams begin testing in an identical, unstructured manner where an agent checks an output and declares it valid without any underlying baseline or accountability. Without counting past failures, a team has no accurate measure of its accumulated technical debt.

software engineer computer terminal code display office

Stock photo for illustration only, not from the actual event

10sFull regression run duration
50Scenarios across 6 test suites
22Hidden HTTP-500 errors caught

The full regression suite eventually ran in just ten seconds, a timeframe fast enough to be skipped—which is precisely how confidence quietly turns into a falsehood. Speed and safety operate as the exact same lever here; rapid feedback enables bolder correct changes while discouraging incorrect ones.

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

Achieving a 10-second test execution cycle directly aligns with the fast feedback loop principle in modern software engineering. When autonomous agents receive rapid validation, they can safely refactor complex architecture layers without fearing undetected regressions, significantly boosting iteration velocity and overall system resilience.

Framing verification as a ledger rather than a gate changes team dynamics completely. While a gate merely halts bad work, a bridge allows productive work to travel safely, empowering AI agents to refactor code with immediate awareness of potential systemic impact rather than shooting blindly into the dark.

Treating failure data as a growing organizational asset shifts testing from compliance checking to cartography of system vulnerabilities. This historical context prevents repetitive debugging cycles across project iterations.

"A test system that can say 'I don't know' is worth more than one that says 'passed'"

Derek Wang

The fulltest framework integrates three essential mechanisms absent from standard runners: classifying failures before judgment using historical ledgers, growing the ledger organically from real incidents and git commits, and raising the operational baseline automatically so the system cannot silently backslide.

Ultimately, a massive test suite means very little if it measures the wrong metrics or never interacts with actual failure paths. The only metric that truly matters is how many past real incidents the suite would have successfully caught. When applied in practice, a 10-second run scaled into 50 scenarios across six specialized suites, successfully intercepting 22 hidden HTTP-500 errors before release.

Source: Dev.to

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article