Test-First is the clear winner of the pre-AI coding era. But how does AI-assisted development change the balance between Test-First and Test-After? Who is the winner now?
I am a long-time enthusiast and expert in test-first TDD and A-TDD. And I have spent 2+ decades leading organisations and teams through major tech shifts and pioneering new ways of working.
As part of the task of reimagining how we code with AI, I am reflecting with openness and curiosity on my experience with AI-assisted coding (human-driven, agentic, frequently reviewed AI-generated code). This experience comes from creating production-grade software.
Specifically, in this post, I document the pros and cons I have observed in the test-first and test-after approaches when used with AI-assisted coding.
Here, I revisit the traditional benefits of Test-First development and compare them directly with Test-After in AI-Augmented Development.
Let’s start from the end. So far, these are my conclusions following from the Pros and Cons I have experienced in my own context.
Unlike before in manual coding, where test-first is a clear winner, now in mid-2026, I see AI-Test-first and AI-Test-after both working well, with the latter (Test-After) being faster and consuming fewer tokens while achieving the same benefits of Test=First. Each has its pros and cons. Further evolution of LLMs’ capabilities may change this; meanwhile, this is where things stand.
A Sr software engineer experienced in TDD and A-TDD, who naturally adopts the A-TDD and TDD way of thinking about behaviour and code design already when writing each prompt, may benefit from the pros of AI-Test-after while countering the cons with an additional step in which they direct the AI to perform mutation testing, fault injection, or adversarial verification.
A Jr software engineer, and everyone without ingrained TDD and A-TDD habits, may benefit from adopting the AI-Test-first approach. This guides the development along the lines of TDD and A-TDD ways of thinking, improving code design quality (resulting from code generation after tests are generated) and test reliability.
Below is the complete list of pros and cons underpinning these conclusions.
See similar conclusions reported by other practitioners: compared to testing afterwards, a test-first approach increases costs and extends timelines without improving outcome quality.
See, for example, Birgitta Böckeler’s findings, here: https://martinfowler.com/articles/exploring-gen-ai/tdd-in-the-agent-loop.html
Stefan Smith, Head of Technical Excellence at Gousto, who experimented with having the agents follow TDD, found that it doubled costs and increased latency by 2 to 3 times. A positive side effect he noted was the reduced cognitive costs of reviewing and thinking about the decisions and code changes made by agents:
Uberto Barbini, Author of “Process over Magic: Beyond Vibe Coding, on the other hand, is experimenting with making the agent write all the Acceptance tests first, so a big batch ATDD, to verify the model understood the specification, before jumping into the implementation. This is halfway between Test first (for the acceptance tests, all in one block) and Test after (for the unit and integration tests).
Which Pros and Cons have you experienced in your own context?
_______________________________________________________________________________
Here are two concepts that can be used as lenses when analysing the pros and cons of AI Test-first and Test-after:
With AI test-after, during the 1st step of code generation
With AI test-first, during the 2nd step of code generation, which follows the tests’ generation
Running mutation testing tools like Stryker or PITest across a full codebase and test suite in CI/CD can take up to 30 minutes.
Here, the goal and granularity are different.
In manual TDD, short Red-Green-Refactor cycles, writing a failing test first, then writing only the minimal code needed to make it pass, prevent false positives and guarantee the coverage of every execution path. That’s the safety net I want to preserve.
Because the AI generates code and tests simultaneously, having the AI follow up with mutation testing on that specific changeset is fast and restores that exact safety net.
The AI spins up isolated worktrees and injects bugs across different execution paths. If a test fails to turn red, the AI immediately uses that feedback to fix the test on the spot.
The AI spins up isolated worktrees and injects bugs into different execution paths. If a test fails to turn red, the AI immediately uses that feedback to self-correct and fix the test on the spot.
It’s also possible to instruct the AI to run mutation testing tools like Stryker or PITest directly within the local workflow before CI/CD. By scoping execution to only the newly modified code and test files, runs complete in seconds. If a mutant survives, the AI captures the log and feeds it into its fast agentic self-correcting loop to fix the weak test automatically (as it does for failing builds, lint, or test failures).
This localised mutation testing and the CI/CD codebase mutation testing steps can coexist and support each other. Like running tests locally before a commit and running them as part of the CI/CD pipeline.

See how we can help across the full AI adoption lifecycle.