Back to Blog

I Tried an AI Browser Agent for the Boring Parts of Freelancing

#ai#automation#browser-agent
I Tried an AI Browser Agent for the Boring Parts of Freelancing

There’s a category of freelance work that never shows up in a portfolio but eats a real chunk of every week: filling out invoices manually across a few platforms that each format things differently, checking multiple clients’ analytics dashboards for a monthly report, researching competitors before writing a proposal, and re-typing the same administrative information into different forms. None of it needs my actual engineering skill — it just needs time and patience for clicking. This month I tried offloading some of it to a browser agent — an AI that can actually navigate a website, not just answer questions about what’s on screen. The results weren’t as clean as I’d hoped, but they were useful enough that I’m sticking with it.

The easy wins: searching and summarizing

What worked most consistently was something like “open these five competitor sites, note their pricing tiers, build a summary table.” That’s a task genuinely suited to an agent — reading a bunch of pages, extracting structured information, arranging it into something I can use immediately. The time I usually spend on competitor research before writing a proposal dropped a lot, from about an hour to roughly fifteen minutes reviewing the agent’s output. I also like not having to write a custom automation script for a task I’ll only run once — in the past, a repeated task that needed web navigation meant deciding whether it was worth writing a Playwright script for it. Now, for one-off tasks, I just give it a plain-language instruction.

It’s not always faster, though, especially for simple tasks. For a one-step task I’d normally do quickly by hand — check one number on one dashboard — waiting for the agent to navigate, screenshot, and verify sometimes takes longer than just opening the tab myself. The actual break-even point is on tasks with many repetitive steps across multiple pages, not simple single-step ones.

Where it failed most often: anything requiring login and sensitive data

When I tried having it fill an invoice on a freelance platform that required login, it could navigate to the right page fine, but I still typed in credentials and bank account numbers myself rather than letting the agent handle those fields. Not because it’s technically incapable — but because I’m not comfortable giving full access to a financial account to a system that can still misinterpret an instruction. That’s a limit I set for myself, not a limit of the tool.

One incident made me a lot more cautious. I asked it to check a payment status on one dashboard, and there were two buttons sitting close together — one “View Details,” one “Cancel” — positioned close because that dashboard’s UI just isn’t laid out cleanly. The agent started to click the wrong one before I stopped it mid-action while watching the live preview. Nothing actually broke, because I was watching in real time, but it was a hard reminder for why I don’t let it run unsupervised on anything with real consequences for a wrong click.

I became a lot more aware of the difference between “reading a page” and “understanding what the page means.” The agent is genuinely good at reading text on screen. But a few times it misread the intent of a visual layout that’s obvious to a human — a red badge meaning “action required,” for instance, where it didn’t register the urgency because the text itself was neutral. Interfaces built for humans often lean on visual signals that aren’t consistently encoded in the raw text, and that’s something the agent doesn’t reliably pick up on yet.

For client reports, a first draft — not the final version

I have it open each client’s analytics dashboard, pull the key numbers, and assemble a draft report. The draft’s a solid starting point, but I always reread it and add context the agent doesn’t have — why traffic dipped this week (scheduled maintenance, not a real problem), or why one number looks off (a tracking typo I already know about but haven’t fixed yet). Without that context, the agent’s report reads as alarming when it’s actually fine.

Privacy and cost, two things I underestimated going in

A browser agent can technically see everything on screen while it’s running — other open tabs, session cookies, whatever’s visible. Before connecting it to the same browser I use to log into client accounts, I read up on how that data gets processed, whether it’s stored, or whether it’s used once and discarded. For anything touching client accounts, I now use a separate browser profile dedicated to the agent, so there’s a clean separation from my own sensitive accounts. It’s an extra setup step, but it felt necessary once I actually thought through what the agent can see while it’s running.

A few multi-step navigation tasks turned out to consume more resources than I expected going in, since every step requires the agent to re-read the current page state. For something run occasionally, that’s a non-issue. But for a task I wanted to turn into a weekly routine, I started actually doing the math on whether the cost was worth it against the time saved, and the answer wasn’t always yes for every kind of task.

My takeaway: this is a useful tool for a specific category of work, not a general replacement for admin tasks. For research and summarizing across multiple sources, I now genuinely use it as part of my routine. For anything touching financial data or credentials, I still keep manual control myself. And for simple one-step tasks, I still just do them by hand because the overhead isn’t worth it. It’s not a magic fix that erases boring work, but it’s meaningfully cut down the share of my time spent on tasks that don’t need any of my actual engineering skill.