I have a full-time job. I also run a small side business, and nearly all of the work in it is done by AI agents. My part is what needs a human: taste, approvals, payments, signatures, and the final click on anything that leaves the building.

I will say the awkward part first. The business is not profitable yet. This essay is about the setup, not results. The hard part turned out not to be the AI. It was making it trustworthy.

The problem: a chat box forgets

Most people use AI as a chat box. You open it, explain your situation, get a good answer, and close it. Tomorrow you explain everything again.

For a side business that is a real problem. The bottleneck is rarely ideas or code. It is memory, follow-through and coordination: what did I decide last month, what is due this week, what was I in the middle of? A chat box has no answer to any of that. Nothing runs while you sleep. Nothing is checked unless you remember to ask.

So I stopped treating the AI as a tool I consult and started treating it as a small team I manage.

What the setup is

There are four moving parts.

Two conductors. These are two independent AI agents that sit in the lead seats. On a hard question, each one answers alone first. I call these sealed answers: neither sees the other's reply. Then they compare. Where they disagree, the disagreement is written down together with the test that would settle it.

Workers in lanes. A worker agent gets one bounded job and writes code in an isolated copy of the project. It cannot touch the real thing. The conductor runs the tests before anything ships.

Memory that lives in files. There is a decision log of about 1,150 lines, and roughly 37,000 receipt files behind it. A receipt is the saved evidence for a claim. If an agent says something is true, it should point to one.

A clock. There are 365 scheduled checks, which I call agenda cues. They fire on their own: recurring admin deadlines, a monthly review of how the whole thing is going, routine system checks.

The rules are few and strict. Tell the truth. Mark inferences as inferences. Never spend money or send anything without my explicit yes. Keep private data in a private folder.

Three things it does well

1. It remembers what I already decided. The log is append-only. A line is never edited; a reversal is a new line that names the old one. The rule says every decision should also state the condition under which it stops applying. In practice only a few dozen lines do so far, which is its own small lesson about rules nobody checks.

Here is the kind of thing that happens. A fresh agent session suggests a pricing option. Before offering it, the agent searches the log, finds that I settled the question weeks ago, and says so: here is the line, and nothing it depended on has changed. I do not re-argue my own decisions at 11 p.m.

2. It acts on a schedule without being asked. A dated reminder I would have forgotten simply shows up, with the relevant numbers already pulled in.

The monthly review is the best example. On the day, a cue fires and an agent assembles what happened, what was spent, what slipped, and what the evidence says about the plan. I read it, push back, and decide. I never had to remember to start it.

3. It builds in bounded lanes. Say a purchase screen has a bug. A worker is given only that bug. It reproduces the failure in its isolated copy, writes a test that fails, then fixes the code until the test passes. The conductor does not take the worker's word for it. It runs the test itself, and runs it against the old code to confirm it really failed before. Only then does the change go forward, and once the checks pass, the agents ship it to phones themselves, in stages, ready to stop at the first sign of trouble.

What goes wrong

Agents are confidently wrong. Any setup that ignores this will eventually embarrass you.

It happens in ordinary ways. An agent reads an outdated copy of the code and reports on a setting that no longer exists. A counting query silently misses a group, and the number still looks reasonable. All of it is delivered with total confidence.

Three habits keep this manageable.

First, two seats. When two agents answer alone and then compare, a shared blind spot is less likely, and a disagreement is visible instead of buried.

Second, executed evidence. A claim counts only if someone ran something and looked at the result. Reading code and concluding it works is a guess, and it gets labelled as one.

Third, written retractions. When an agent finds that an earlier claim of its own was wrong, it says so in the log, plainly, and says what replaces it. There are 43 of these so far.

I used to read that number as a bad sign. I now read it as the system working. A team that never corrects itself is not a team that is never wrong. It is a team that is not checking.

What it costs in human time

My target is three to five hours a week. It is a target, not a measured result, and a busy week breaks it.

Those hours go to the things an agent should not do for me. Judging whether something looks and feels right. Approving or declining a proposal. Making a payment. Signing. Pressing send. I also read the retractions, because they tell me where the system is weakest.

There is a money cost too, for the model usage that powers all this. I am not giving a figure here.

When the machine does the work, choosing is the job

A friend watched this setup run one evening and asked me something I keep coming back to: at what point is it no longer me, but the AI living my life? It is the old puzzle of the ship whose planks are replaced one by one until none of the originals are left. Is it still the same ship?

My answer, for now: the planks are the chores. What makes it mine is what I choose. Which day I register the business, which offer to make, what to refuse. The agents can swap out every task. They do not get to pick the course.

This is not only a solo-builder question. In early October, OpenAI published more than 700 mathematics papers produced by an internal model, many with proofs checked by computer. Mathematicians are still working through them, and some called the release a show of force rather than scholarship. Whatever the verdict, the direction is clear. When a machine can produce the work, the scarce human jobs are choosing which problems matter, checking that the result is right, and taking responsibility for what gets shipped.

That is also why my system is built the way it is. The agents produce. I prioritise, I approve, and anything with my name or my money on it stays with me. As the output gets cheaper, the choosing and the checking get more valuable, not less.

Honest status

Here is where things stand. The system runs daily and I rely on it. The side business it serves is not profitable so far.

That is not a pitch. It is the reason I am writing. I would rather show the working parts, including the retractions, and let you judge whether the approach holds up as results arrive.

From here I plan to build in public: what the agents get right, what they get wrong, and what I change as a result. If the system cannot make the business profitable, you will read that here too.

If you want to follow along

If you are a solo builder juggling a job and something of your own, and you want to see how this develops, join the waitlist: [email protected]