A Month With an AI Code Review Bot on My PRs: What Actually Stuck

Code review on my solo projects has always had exactly two stages: I write it (or read what an agent produced), then I read back over it myself before merging. No real second pair of eyes, because, well, I work alone. This month I tried closing that gap by putting an automated AI review bot on a few repos — the kind that comments directly on a PR every time you push, like a human reviewer but on a timescale of minutes. I’m not writing this to push a specific tool. Just being honest about what actually stuck and what became a notification I swiped away.
Setup was trivially easy. Connect it to GitHub, grant repo access, and within minutes of opening a PR there’s an automated comment breaking down the changes file by file. Zero friction to start trying it. What I didn’t expect was how quickly I started being able to predict what kind of comment was coming before I’d even read the diff myself.
What actually helped: catching the boring stuff that’s easy to miss when I’m tired. An unused declared variable, an unnecessary import, small inconsistencies like one function using async/await while the one next to it still uses a .then() chain when they should match. Exactly the class of problem I’d normally catch myself if I were fresh, but the kind of thing that slips through at 11pm racing a client deadline. The bot never gets tired, and that turned out to be genuinely valuable at this level of feedback.
What didn’t help: high-level architecture suggestions that sound smart but don’t understand the project’s context. A few times it suggested refactoring toward a “more best-practice” pattern that didn’t fit a constraint it simply couldn’t know about — a client project deliberately kept simple because it’s small in scope and will be handed off to a junior internal team, where over-engineering is actively counterproductive. The bot has no access to context like that, and there’s no easy way to teach it besides manually dismissing the comment every time.
False positives showed up often enough that I nearly gave up in week two. A few times it flagged a “potential null pointer” on code that was already guarded a few lines earlier, just through a helper function it didn’t trace through. If I skim comments without thinking, I can waste time double-checking something already safe. That’s what made me realize an AI review bot isn’t that different from a smarter linter — useful if I treat it as a signal to check, not a verdict to trust outright.
The most useful thing wasn’t bug-catching at all, it was summarizing large diffs. When I’m working with a coding agent on a fairly big change spanning a dozen-plus files, the review bot’s narrative summary — “this change refactors input validation across three endpoints and adds one new middleware” — gets me oriented faster before I read line by line myself. It doesn’t reduce how much I actually read, I still read all of it. It just cuts down the mental “loading” time before real review starts.
I turned off comments on a few file types, and that cleaned up the signal considerably. At first it was also commenting on config files, auto-generated SQL migrations, and test snapshots — noise no human needs reviewed at all. After excluding those, the ratio of genuinely useful comments went up a lot. Might be obvious to other people, but it only clicked for me after living with it: a tool like this is only as good as how it’s configured, not just how good the model behind it is.
I haven’t turned this on by default for client projects. Not because it’s useless there, but because some clients have policies about third-party tools touching their code, and I’m not going to assume that’s fine without asking first. That’s the kind of consideration that never shows up in articles that just cover “how to set up an AI reviewer” — in freelance work there’s a layer of permission and client trust that has to clear before any tool, no matter how good, gets installed.
One moment genuinely changed my view of this. Week three, it caught an API key I’d accidentally hardcoded. I was rushing through debugging a third-party API integration, dropped a test key directly into the code to test something quickly, meant to remove it before committing. Forgot. The review bot flagged it within minutes of the push, well before I’d have noticed myself. Fortunately it was still on an unmerged branch and the key was a scoped test key, so the actual damage was minimal. But that moment made it clear the value here isn’t just style consistency — there’s a class of mistake that’s genuinely expensive if it slips through, and it’s exactly the kind of mistake I’m worst at catching myself, because I already know it’s my own code and I’m automatically less skeptical rereading it.
I also started comparing its notes against mistakes I remember making myself in the past — forgetting to await a promise inside a loop and creating a subtle race condition, for instance. The bot consistently catches patterns like that, which gives me an extra layer of confidence specifically for the cases I already know are risky but sometimes slip through when I’m tired.
The verdict after a month: I’m keeping it, but not as a replacement for my own review. It functions more like an assistant that catches small things I miss when I’m tired, and gives me a fast orientation on big diffs — not a second reviewer that actually understands the project. I still read every diff line by line before merging, same as before. There’s just one extra layer now that catches a handful of small things before they reach my own eyes.