Back to Blog

A Month of Daily Coding With Claude Sonnet 5

#ai#claude-sonnet#workflow
A Month of Daily Coding With Claude Sonnet 5

I switched my default coding agent model over to Claude Sonnet 5 about a month ago, mostly because it landed as the cheaper option for running agents and I wanted to see if “cheaper” meant “worse” for the kind of work I actually do. I’ve been running it across three active freelance projects since — a Next.js dashboard, an Avalonia desktop tool, and this site — so this isn’t a benchmark post, it’s just what changed for me in practice.

The first thing I noticed wasn’t speed, it was consistency. With the model I was using before, I’d get long stretches of genuinely good output punctuated by weird misses — renaming a variable halfway through a file and not catching it, or forgetting a constraint I’d stated three messages earlier. Sonnet 5 doesn’t feel dramatically smarter on any single task, but it drops that context noticeably less across a long session. On the dashboard project, I had a session that touched eleven files over what must have been forty-plus tool calls, and it still respected a naming convention I’d mentioned once at the very start. That’s the kind of thing that used to require me re-stating constraints every few turns, which ate more of my attention than the actual coding did.

Cost changed how I use it, not just what it costs. Because it’s cheaper per token, I stopped being stingy about giving it context up front. Before, I’d trim down what I fed it — just the relevant function, not the whole file, not the related test — partly to save cost, partly because more context sometimes meant more noise. Now I just point it at the file and the two or three related files and let it read. The result is fewer “wait, that’s not how we do it here” corrections later, because it actually saw the existing pattern instead of guessing at one.

Where it’s genuinely better: multi-file refactors. I had it move a chunk of shared validation logic out of six duplicated API routes into one shared module — exactly the kind of mess I complained about seeing in inherited codebases. The old workflow for this used to be: agent does one file, I check it, agent does the next file, repeat, catch drift between files at the end. With Sonnet 5 I described the target shape once and let it work across all six routes in one pass, and the drift between files was close to zero — same function signature, same error handling pattern, same import style, in every file. I still read every diff line by line before committing, because that discipline doesn’t change just because the model got better. But there was a lot less to fix.

Where it’s not magically better: judgment calls. Anything that requires an actual product or architecture decision — should this be a new table or a JSON column, should this state live in a store or stay local to the component — it still just picks something reasonable and moves on unless I stop it and ask it to lay out the tradeoffs explicitly. Fine, and honestly the correct default behavior for a tool, but it means the parts of my job that were never going away — deciding, not just producing — still haven’t gone anywhere. If anything, the fact that execution got faster and cheaper makes the decision-making part a slightly larger share of what I spend my time on now.

The desktop app project surprised me the most. Avalonia/C# is a smaller niche than web frameworks, which usually means worse model performance because there’s less training data to draw from. I expected more hand-holding here. Instead it handled XAML binding patterns and the MVVM structure I’d already established about as well as it handled the Next.js routes — it clearly picked up the existing conventions from reading the codebase rather than falling back on generic C# patterns that didn’t match how the project was actually structured. That’s the part that matters most for a solo dev working across a genuinely mixed stack: I don’t want a tool that’s great at the popular framework and mediocre everywhere else, because my actual week doesn’t look like that.

The thing that took me longest to adjust wasn’t the model, it was my own habits. I spent the first week or so still working in the smaller, cautious chunks I’d built as a habit around the old model’s context slips — checking in after every file, re-stating things I didn’t actually need to anymore. It took a deliberate effort to notice I was compensating for a limitation that wasn’t really there anymore. A strange kind of lag — not the tool being behind, but me being behind the tool — and I suspect it’s a bigger factor in how much value people actually get out of a model upgrade than the benchmarks suggest.

Would I recommend switching? If you’re already running an agentic coding workflow and cost was ever a factor in how carefully you scoped your context, yes — the cheaper-per-token part changes your behavior in a good way, not just your bill. If you’re expecting a model swap alone to fix a codebase with no review discipline behind it, it won’t — the same point I keep landing on with every AI tooling post I write this year: the tool got better, the job of actually reading what it produces did not get any smaller.