
This idea of running my entire freelance business on AI started as a bet with myself.
I had been writing about AI tools for months. Recommending them, explaining them and breaking down which ones were worth paying for and which ones were not. But I was using them selectively, the way most freelancers do. A bit of Claude for drafts, some Otter.ai for meeting notes and Canva for the occasional image. I was not running my business on AI, I was just dipping into it occasionally and calling it a workflow.
So I decided to find out what it actually looked like to go all the way. Thirty days, every business task, AI assisted or AI led wherever I could make it work. Client proposals, content delivery, invoicing follow-up, research, social media, scheduling, financial tracking, all of it. I kept notes every day, some days the notes were excited while in some days they were not.
The Setup
I did not stop thinking neither did I hand over client relationships to a chatbot. What I committed to was reaching for an AI tool as my first move on every business task rather than my second or third.
The tools I used throughout the thirty days:
Claude for all writing, drafting, editing, client communication and strategic thinking tasks. ChatGPT as a backup when I hit Claude’s daily limit. Otter.ai for every client call and Fathom as a second layer on some calls. If you want a deeper look at what each of these tools actually does before committing to a month with them, our breakdown of the AI tools freelancers are using to increase output covers them in detail.
Perplexity for research while Canva with its AI features for images and graphics. Notion AI for organizing my work and client notes. Wave for invoicing and expense tracking with AI-assisted categorisation. Buffer for social scheduling while Fireflies is for post-call intelligence on two calls where I wanted the talk-time data.
Total monthly cost at the time of the experiment was $47. Claude Pro at $20, Otter.ai Pro at $16.99, and Fathom Pro at around $10. Everything else was on free tiers.
For freelancers who cannot justify that spend yet, we ran a separate experiment on whether free AI tiers are enough to run a freelance business. The answer is more nuanced than most people expect.
I had five active clients going into the experiment. Two were ongoing retainer relationships, two were mid-project. One was new, just signed the week before the experiment started.
Week One
I will be honest with you, the first week felt like cheating.
The first thing I noticed was how much time I was saving on proposal writing. I had an inbound inquiry on day two from a potential client who needed content strategy work. Before the experiment, a proposal like this would take me about ninety minutes. I would draft it, revise it, second guess the tone, rewrite the opening and spend the last twenty minutes on pricing language.
With Claude, I wrote a rough version of the brief in about eight minutes, describing the client’s situation, what I understood their problem to be and what I would offer. Claude turned that into a structured proposal draft. I spent thirty minutes editing it toward my voice and adding the specific details only I would know about the client’s market. Total time spent was thirty-eight minutes. The proposal was better than my average because Claude pushed me to address objections I would typically leave unspoken.
The client signed within two days.
By day five I had processed three client call transcripts through Otter.ai. Written four content pieces with Claude assistance, responded to eleven client emails using AI-drafted versions that I edited and updated my financial tracking in Wave. I had also posted six times across my social accounts using Buffer with copy drafted by ChatGPT.
My total billable hours for the week were eleven. My total hours worked, including all the business tasks around the billable work, were fourteen. That ratio felt unusually good. In a typical week without this level of AI integration, the same eleven billable hours would have come with closer to nineteen or twenty hours of surrounding business work.
Three hours saved in a week felt meaningful. I was optimistic going into week two in a way that I should probably have been more cautious about.
Week Two

The first crack appeared on day nine. I had a client call with one of my retainer clients, a marketing director at a UK-based software company. We talked for forty minutes. Otter.ai transcribed the whole thing and produced a clean summary. I read the summary and felt like I understood the call.
Three days later she emailed asking about something specific we had discussed, a concern she had raised about how a particular piece of content might land with a segment of their audience that had complained about similar messaging in the past. It was a nuanced concern, mentioned briefly, and I had registered it in the moment but the Fathom summary had reduced it to a single line about “audience sensitivity considerations.”
I went back to the full Otter transcript. The concern was there, more detailed and more specific than the summary had captured. But I had read the summary rather than the full transcript because that was the efficient thing to do. I had let AI compress a piece of information that required human attention and I had not caught the compression until it mattered.
The client was fine, I addressed it in the next piece. But the incident made me reconsider how I was using the transcription tools. The summaries were useful for recalling what had been covered. They were not reliable for capturing what had been felt, what had been said tentatively, what the client had seemed uncertain about. That layer required reading the full transcript. Or better paying careful attention during the call rather than relying on the post-call tools to reconstruct what you should have been present for.
I adjusted from day ten onwards I used Fathom summaries as a starting point. And always cross-referenced with the full Otter transcript for anything client facing.
The second crack appeared on day eleven. I had been using Claude to help me draft my client emails throughout the first week and the early part of week two. The drafts were clean, professional and well structured. They were also starting to sound the same.
I noticed it because one client replied to an email with “this feels a bit formal, is everything okay?” She had been working with me for eight months and she knew my natural communication style. The email I had sent was not rude or cold, it just did not sound like me. It sounded like Claude. That moment connected to something broader about what AI writing tools do to your voice over time. Covered honestly in a separate article that this experiment ended up confirming.
I went back through the emails I had sent that week. Four of them had the same structural pattern: opening acknowledgment, main point, secondary point, closing offer of availability. They are clean, professional and have Identical rhythm.
From day twelve I started writing the first sentence of every client email myself before handing it to Claude. That one sentence, written in my actual voice, changed the quality of the output. Enough that two more clients commented positively on my communication that week without knowing I had changed anything.
Week Three
By week three something like a working rhythm had settled in. The tools stopped feeling like a conscious choice on every task and started sitting more naturally in the background of how I was working.
The AI tasks that were working consistently:
Research was genuinely faster. For a content piece on regulatory changes in fintech, I used Perplexity to surface recent developments. It pulls from live sources rather than a training cutoff, which matters for regulatory content. Then asked Claude to help me understand the implications for a specific audience segment. What would have taken me two hours of reading and synthesis took forty-five minutes. The output was accurate. I fact checked three specific claims independently and all three held up.
First drafts were consistently good enough to edit rather than rewrite. This sounds like a small distinction but it is significant. A first draft that needs to be rewritten is demoralizing and time consuming. A first draft that needs to be shaped is energizing. Claude’s drafts were almost always in the second category when I gave it enough context.
Financial tracking was genuinely painless for the first time in my freelancing career. I had always hated updating Wave. With Notion AI helping me organize my weekly review and Wave’s categorisation handling the mechanical parts. I was spending about fifteen minutes per week on financial admin rather than the forty-five it typically consumed.
Scheduling and social content felt almost fully automated in a way that was comfortable rather than hollow. The posts did not feel like me at their most interesting, but they were competent and consistent. Which was better than what my social presence had been before the experiment, which was infrequent and inconsistent.
The AI tasks that were not working or working less well than expected:
Strategic thinking was the area where I pushed AI hardest and found the most resistance. I was working on a positioning proposal for the new client, trying to articulate a strategic recommendation about how they should approach a specific market segment. I tried to use Claude as a thinking partner and gave it context. I asked it questions and it produced options. None of the options felt like the right one and I could not quite explain why.
I spent an hour going back and forth with Claude trying to get to the thing I could feel was the right answer but could not articulate. Eventually I closed the laptop, went for a walk and came back with the positioning in my head fully formed. I opened Claude, described it and Claude helped me write it up clearly.
The thinking had happened away from the tools not with them. That felt important.
Client relationship management was the area I was most careful about and where AI helped the least proportionally. I had a difficult moment in week three where one of my mid-project clients pushed back on a deliverable in a way that felt personal. The email was curt, feedback was vague and frustrated. I drafted a response with Claude’s help, read it back and deleted it. The response was technically correct and professionally phrased and completely wrong for the moment. I wrote my own response and it was messier and less polished and almost certainly more effective.
Some emails should not be optimized for professionalism. They should be optimized for being a human being in a professional context, which is a different thing.
Week Four: The Honest Accounting
By week four I was tired in a specific way I had not anticipated but not from overwork.
My hours had actually gone down. I was tired from the constant context switching between my own thinking and the tool’s output. Every task had an additional cognitive step: give the AI enough context, evaluate what came back, decide what to keep and what to change and then edit toward my voice. That sequence was efficient compared to doing everything from scratch. It was also more mentally present than my old workflow, which involved a lot of mindless execution that the AI had now replaced with mindful supervision.
I talked to two other freelancers during week four who had done similar experiments and both described the same thing. One called it “AI overhead,” the thinking required to use AI well rather than just using it. The other said she felt like a manager who had delegated everything and then had to review everything, which is a different kind of busy than just doing the thing yourself.
On day twenty-six I made a deliberate decision to write one piece entirely from scratch with no AI assistance. A reflection piece for my own newsletter but it took longer and the first draft was rougher. And at the end of it I felt something I had not felt in four weeks, the specific satisfaction of having made something entirely by myself.
I am not sure what to do with that feeling analytically. I note it because I think it is real and worth naming.
The Numbers at the End of Thirty Days

Revenue was $4,840. My average month in the three months before the experiment was $3,920. The increase was partly the experiment and partly the new client who came in at a good rate. I cannot cleanly attribute the difference to AI alone.
Billable hours worked was 67. My average in the three months before was 74 billable hours per month. So fewer hours more revenue, the math looks good.
Total business hours including non-billable work was 89. My average before was 112, that 23 hours reduction is the number I find most significant. Twenty-three hours over a month is almost three full working days. That time went into things I had been meaning to do for months, rest, a side project, two long overdue conversations with people I had been too busy to call.
Proposals sent was 8 and 3 out of 8 was signed. My usual conversion rate on proposals is around 25 to 30 percent. This month was 37.5 percent. Sample size is too small to draw conclusions but the direction is positive.
Client complaints or negative feedback was just 1, the email that sounded like Claude rather than me. Was resolved with no lasting damage.
Tools I stopped using mid-experiment was 0. Every tool I started with I continued using.
Tools I added mid-experiment was 1. I added Loom on day eighteen for client deliverable walkthroughs after realizing I was spending too long writing explanatory emails that a two-minute video would replace more effectively.
What I Kept, What I Changed, What I Dropped

At the end of thirty days I did not go back to how I was working before. But I also did not keep everything.
What I kept permanently:
I kept Otter.ai for every client call, with the full transcript reviewed for anything that felt significant rather than just the summary.
Claude for first drafts of almost everything, with a personal opening line written by me first for client-facing communications.
Notion AI for my weekly review and client workspace organization.
Wave with AI-assisted categorisation for financial tracking.
Buffer for social scheduling with AI-drafted copy that I edit to add one specific, personal detail per post.
What I changed:
I stopped using Claude for strategic thinking tasks not because it is bad at it. Because I found that my thinking was better when I arrived at the AI with a half formed idea rather than asking the AI to do the forming. The tool is better at articulating and structuring ideas I already have than at generating the ideas in the first place. At least for my kind of work.
I stopped letting AI draft difficult or emotionally sensitive client communications. Those I write myself.
What I dropped entirely after the experiment:
I dropped Fireflies. The conversation intelligence layer was interesting but added a step I found I was not using productively. Fathom plus Otter was sufficient and adding a third transcription layer created more to review rather than more clarity.
What I Actually Found Out
Thirty days of running everything through AI did not make me more productive in the sense of producing more output. It made me more productive in the specific sense of producing the same or better output in less time. Which freed up space for the parts of the work that genuinely could not be delegated.
Those parts were mostly the things that required me to be a specific person who had been in a specific room with a specific client. Strategic thinking that only worked after I stopped trying to do it. An email that needed to sound like someone who had been in that working relationship for eight months. A positioning recommendation that came from having seen similar problems fail in similar ways before.
There is a real distinction between the work that can be fully described before you start it and the work that only becomes clear as you do it. AI handles the first category well and the second category poorly. I knew this in a general sense before the experiment. The thirty days made it specific enough to be useful. I now have a clearer sense of which tasks in my particular work belong to which category and I route them accordingly.
Most people using AI tools extensively probably do not have that map yet. Running the experiment gave me mine, which is probably the most durable thing I got from it.
Frequently Asked Questions
Could a newer freelancer with fewer established clients run this same experiment?
Yes but the results would look different. The parts of the experiment that saved the most time, research, drafting, financial tracking are available to any freelancer regardless of experience level. The parts that required personal judgment, difficult client communications, strategic positioning, would have a different shape earlier in a freelancing career because the judgment itself is less developed. A newer freelancer might find AI fills more gaps initially, since the judgment-requiring tasks come up less often. The risk is different too: without established client relationships to catch the moments when AI makes you sound like AI, the feedback loop is slower.
What did the AI tools miss that you noticed only in hindsight?
Two things specifically. The client concern about audience sensitivity that got compressed in the Fathom summary was one. The other was subtler: across the month, my Upwork profile activity and a couple of industry forum responses were drafted with AI help and I noticed only at the end of the month that my public voice had drifted toward something more generic than my natural writing. Nobody commented on it negatively. But when I went back and read several months of my own forum posts in sequence, the AI assisted ones were identifiable. Not bad just slightly less specifically mine.
Was the revenue increase meaningful or just noise?
Honest answer is partly noise. The new client contributed significantly to the monthly total and that client came in through a referral that had nothing to do with AI. The improvement in proposal quality probably contributed to the 37.5 percent conversion rate but the sample size of eight proposals is too small to draw a real conclusion. What I can say with more confidence is that the 23 hour reduction in total working hours was not noise. That pattern held consistently across all four weeks and the work quality did not deteriorate.
What would you do differently if you ran the experiment again?
I would define upfront which tasks I was keeping fully human throughout the experiment rather than discovering the hard way which ones needed to stay that way. The client email that sounded like Claude and the strategic thinking session that only resolved after a walk were both things I could have predicted would not benefit from full AI delegation if I had been more deliberate at the outset. Running a cleaner experiment with intentional human-only zones would produce more useful data.
Is $47 per month a realistic AI budget for a freelancer starting out?
The $47 covered three paid tools, the rest of what I used was on free tiers. For a freelancer in the early months who cannot justify that spend yet, the free tiers of Claude and ChatGPT alternating, Otter.ai free at 300 minutes per month and Notion free for organisation would replicate most of what the paid setup delivered. The main thing the paid tiers added was removing the daily and monthly limits, which mattered in a high activity month like this experiment. In a lighter month, the free tiers would have been sufficient for most of what I did.

Johnson Alaekezie is a freelancer and digital content creator with several years of experience working across writing, content strategy, and digital services. He founded IncomeGigAI to share honest, practical information about building income online without the hype that dominates most of this space. Johnson has worked with clients in the US, UK and beyond, and writes from direct experience rather than theory. He is based in Nigeria and covers the digital income topics he has personally navigated as a working freelancer.
