Back to Blog

Using AI to Write Tests: What Helps and What I Still Check Manually

#ai#testing#productivity
Using AI to Write Tests: What Helps and What I Still Check Manually

Writing tests is the kind of work I keep putting off. Not because it’s hard — just repetitive and time-consuming, and whenever a deadline’s close, tests are always the easiest thing to cut first. AI coding agents turned out to be a great fit for covering the most tedious part of this, but after a few months of actually relying on it, I’ve also gotten a lot clearer on where the limits are.

What actually helps

Generating test cases for clear scenarios is the easiest thing to hand off to AI — valid input, invalid input, common edge cases like empty strings or empty arrays. I give it one function and ask it to cover every reasonable input combination, and the result is usually 80% done without me having to think it through from scratch. Setup/teardown boilerplate is easy too — the pattern’s usually already obvious from other tests in the same file, so AI just has to copy what’s already there.

Concrete example: last week I added input validation to a form on a client project, and it only took two or three prompts to get test coverage that would’ve normally taken me half an hour to write by hand.

What I still check manually

This part matters more than it looks: whether a test actually tests something meaningful, or just covers a line of code without a real assertion. I once found a test that came back “green” with an assertion that was just expect(result).toBeDefined() — technically passing, but it wouldn’t catch it at all if the logic broke completely. That’s a test that’s more dangerous than no test, because it gives whoever’s looking at a fully-green coverage report false confidence.

Domain-specific edge cases are still something I have to think through myself. AI doesn’t know what bug caused problems in production last month — it won’t automatically write a regression test for that specific case unless I explicitly ask. I’ve gotten into a habit now: every time a production bug gets fixed, I ask AI to generate the regression test for it, while explaining exactly why the bug happened in the first place.

Reviewing AI-generated tests matters as much as reviewing production code

This is the part I was slowest to learn. Early on I’d skim generated tests — see everything “passed,” commit. Now I read every assertion one by one, with the same discipline I’d apply to a production code diff. A wrong test isn’t just useless, it’s actively dangerous — it makes me trust that something’s covered when it isn’t.

The result now: I write tests far faster for the repetitive parts, and the time saved goes into thinking through the scenarios that genuinely need human judgment — things I only know because I understand the business context, not something that can be guessed from the code pattern alone.