First post in a series on keeping software quality high when agents write the code.
For most of software engineering's history, writing the code was slow and expensive work, because trained people had to do it, at human speed and human rates. It is no longer. Agents now produce code faster than any human can read it, and the volume keeps climbing: by mid-2026, 90% of professional developers were using AI coding agents at work at least weekly, 68% daily1 and they report roughly 47% of their code as fully agent-generated against only 27% written fully by hand.2 Telemetry from 22,000 developers shows the same surge in raw output,3 with average pull-request size up 51.3% and the files each developer touches per month up 149.9% once most of a team's developers were using AI tools weekly, compared with those same teams before the shift.
Assuring that the code is correct, safe to ship and something the team can live with for years did not keep pace. It got slower. As AI adoption rose, the median time to first review climbed 156.6% and the median time a change spent in review climbed 441.5%.3 Reading, reworking, testing and reviewing still run at human speed. Work that matters as much as ever, because what arrives is »More code. Not production-ready code.«3
So the constraint moved. And resurfaced on the quality assurance side.
The hard part is not that agents write bad code. It is that their bad code and their good code look the same. Mattis Kämmerer, one of our developers at Teamscale, put it plainly: »Claude always sounds plausible. That is the trap.«
DORA has a name for what that resemblance costs: a »verification tax«, the cognitive load of rigorously auditing »AI-generated code that looks remarkably similar to correct code«.4 Nothing on the surface of a change says which kind it is, so every change pays.
Plausible and correct diverge easily when a codebase is large enough to matter. On Teamscale's roughly 2.2 million lines, the agent only ever sees a slice of the system. It takes an early return, concludes it has understood, and produces a confident, well-formatted, well-argued answer that is wrong. Nothing in its manner marks the difference.
When Mattis investigated how Teamscale handles duplicated uploads of analysis findings, metrics data and code coverage, the agent produced a plausible theory for each. All three were wrong. The agent plausibly assumed a single consistent rule, but the three cases do not follow one: coverage is merged, for findings the later upload wins, and for metrics the result turned out to be arbitrary, a genuine bug Mattis filed. What settled each question was having the agent build automated system tests and run them against the real implementation: »It is like an oracle where I can ask the implementation: how do you actually behave?«
Note what did the work there. Not more careful reading, and not a second opinion from another model, but a deterministic check against the real system, which returns the same answer no matter how convincing the story around it sounds.
Those system tests Mattis created are part of the agentic harness: the things a change has to pass that cannot be argued with. Agents include some such checks into their harness on their own: they compile, run the tests, and read the errors back. Other well-established checks need to be added deliberately: whether a change is actually tested, whether it respects the architecture, whether it repeats a known defect pattern, whether it duplicates something that already exists. These have been standard practice for years, and each delivers deterministic answers.
What changed is what they are worth. Every question the harness settles is a question that no longer has to be paid for in human attention, and human attention is exactly what the verification tax is levied on.
The bottleneck moved from writing code to trusting it. Generating faster does nothing for that, and neither does reading harder. What buys trust back is a harness: long-standing checks, extended deliberately, running at the rate the code arrives so the work does not pile up. That is work we do at Teamscale, and why we think it matters even more now.
This post names the problem. In the following posts, we will explore what teams can actually do about it: what a deterministic harness can settle inside the coding session, before a person ever reads the change; what is left for human review and how to aim scarce attention at it; and what a team has to watch to catch the drift that no single diff shows.
A new post follows every one to two weeks. The Teamscale newsletter carries each one as it appears, for anyone who would rather not check back: sign up below.
1 JetBrains, »AI Coding Agents: Adoption Trends«, Developer Ecosystem Survey 2026.
2 JetBrains, »How Much Code Do Developers Really Let Agents Write?«, Developer Ecosystem Survey 2026.
3 Faros, »AI Engineering Report 2026: The Acceleration Whiplash«.
4 DORA / Google, »The ROI of AI-Assisted Software Development«, v. 2026.1.