

Every “rogue AI” story has a common thread
Ellen Choi, founder of Edgefield Group, built an AI chief of staff she calls TARS.
It mistook a legitimate but unusual purchasing pattern for duplicate payments and recommended auto-refunding thousands of dollars in real revenue. The only reason it didn’t happen is that Choi never gave TARS the authority to execute refunds on its own.
Byron Patrick, Senior Product Manager at Karbon, asked his assistant to summarize his thinking after a customer call. Instead of a summary, it created a shared document and drafted a Slack message to his team.
Nothing left the company this time. But nothing stopped it from trying.
I’m not exempt from this, either. I set up Claude on my work account to draft emails and stage them for me instead of sending them. I never applied that same rule to my personal account. Last week I asked it for help with something, and it just sent an email as me.
No draft. No review.
There was nothing bad in it, thankfully. But it wasn’t really my voice, and I didn’t get a say.
This is a controls problem
We talk about these incidents like they’re bugs in the model’s judgment. They’re not.
Arthur didn’t hallucinate a fake invoice. It made a reasonable-sounding call with real access and no one checking its work before it acted.
TARS pattern-matched on legitimate data and had the authority to act on a bad guess until a human thought to take that authority away.
These stories have the same root cause: nobody separated “figure out what to do” from “actually do it.”
We already solved this problem. Just not in accounting software.
Coding tools figured it out fast. Claude Code and similar tools default to an “ask, then build” pattern.
The agent proposes the plan, you approve it, then it executes.
After his Slack incident, Patrick’s fix was to adopt a personal standing rule that the AI gives him a plan before it acts. That’s the same principle, applied by hand, because the tools don’t enforce it for him.
Accounting software hasn’t caught up.
I see plenty of “draft this” and “post this” prompts in the tools coming to market. I don’t see the deliberate pause between them built into the product.
It’s on the accountant to remember to ask for a plan first, just as it was on me to remember which of my own accounts had guardrails turned on.
That’s backward.
An intern doesn’t get to void an invoice or email a vendor without a second set of eyes, no matter how confident they sound.
We don’t hold AI agents to that same basic discipline.
Your control can’t be hoping the agent behaves
When you evaluate AI tools for your firm or your finance team, the question isn’t “does it make good decisions?”
Every agent looks great until the one time it doesn’t.
The real question is, what happens between the decision and the action? Is there a checkpoint? Can a human see the plan before it executes? Can you revoke authority for one task without turning off the whole tool?
If the answer is no, you have a liability waiting for its moment. It’s like giving an intern the admin password and not reviewing their work.
We’ve spent decades building segregation of duties into processes humans touch. AI agents don’t get a pass just because they’re software.
If anything, they need more separation between judgment and execution, not less. Because it’ll act on a bad guess with total confidence and zero hesitation.