I use AI every day, every single day, and I'm still not the person who can fully trust it. Every now and then I catch a mistake: a wrong assumption, an action that goes against what I asked for, a number that looks right until I check it. That's not going away. AI is probabilistic, so it will keep producing plausible answers that are sometimes wrong.
What can change is how much of your time it costs you. I'll be walking through my full framework for this at the AFP FP&A Series on August 26, three types of AI work, and the check that matches each (by the way, this is a free event, and you can still sign up).
This newsletter is going somewhere the deck can't: two editions, starting today, on the actual techniques and skills you can build in, so checking AI output starts being part of how the work runs.
Today: AI as Automation and Accelerator, and how to verify each.
Next week: AI as a Thought partner, and how to refine it.
Updates and upcoming events:
Claude in Action Cohort, starting September 2, is full. I'm not taking more sign-ups, and I won't run the next one until November.
If your company is looking for a custom program, reply to this email, and I'll get in touch.
I also do 1:1 coaching. If you need a personal session to help you build something in Claude or build the implementation roadmap, reply to this email.
Upcoming events: AFP FP&A Series on August 26th. Free, online.
I will also present at several events in Houston and Austin this fall—stay tuned for updates.
The Check That Matches the Work
Most people check AI output the same way, no matter what produced it. Skim it, see if anything looks obviously off, move on. That holds up for some AI work and fails badly for the rest, because the way it fails is different depending on what kind of work AI is actually doing. A rule that quietly stops matching the business does not look wrong. A fluent paragraph built around a fabricated number does not look wrong either. One review habit applied everywhere misses both, and it misses them in different ways.
Three types of AI work, and each one fails on its own terms

Type | What it means | Shows up as |
|---|---|---|
Automation | AI runs a process for you, start to finish, on the same trigger every time | Reconciliations, mappings, threshold flags, scheduled loads |
Accelerator | AI drafts something, a person reads it and adjusts it | Variance write-ups, board narratives, scenario write-ups |
Thought partner | AI is there while you think; nothing ships without you | Pressure-testing assumptions, structuring an argument before you commit |
I won't walk through the full checklist for each type today; that's a longer conversation for another time.
What matters here is one idea: the type of work sets the standard for the check.
In Automation, you check the rule.
In Accelerator, you check the output against the specific places it tends to break.
In Thought partner, you challenge the assumptions and look for the angle you haven't considered yet, before you land on a conclusion.
This can sound like checking AI output ends up costing more time than just doing the task yourself. It doesn't have to.
Build the check into the workflow so it stops being a step you remember to do and becomes part of the process.
Automation: check the rule, not the output
Workflow | What goes wrong | Check built in |
|---|---|---|
Reconciliation (bank, account, subledger to GL) | A mapping rule drifts, an account stops getting picked up, totals still print clean | Total ties to a control figure. Nothing exceeds a set tolerance. Every line matches a known account, or routes to exceptions |
Multi-source calculation or consolidation | Component parts stop summing to the total once one source changes | Component sum equals total on every run. Row counts match expected source counts |
Threshold-based variance flagging | The threshold stops matching the current scale of the business, so real outliers go unflagged | Flag rate compared against an expected range each cycle. A sample of unflagged items spot-checked on rotation |
Scheduled data load or actuals refresh | The source format changes upstream, rows get dropped or miscategorized without anyone noticing | Row count reconciled against the source file. Unknown categories route to exceptions, never dropped silently |
Cost or overhead allocation | The allocation base shifts, but the rule keeps splitting costs the old way | Allocated total ties to the pre-allocation total. Allocation base reconciled to current org structure each run |
Accelerator: check the output against where it tends to break
Check | What it means | What it catches |
|---|---|---|
Footing | Do the numbers tie across the document | A total that doesn't match its parts, a figure quoted twice with different values |
Causality | Does the driver actually explain the movement | A narrative that sounds plausible but doesn't match what moved the number |
Data integrity | Is every claim traced to a source you opened yourself | A benchmark or citation that was assumed, not verified |
There's a fourth check people I usually list here, accountability: would you defend this as your own. It's left off this table on purpose. It isn't something you check, it's the standard you hold the other three to.
Both tables show what to check. What they don't show is how to actually wire that into a live workflow, or how to run it without doing it by hand every time.
That's the subscriber section: the build for embedding the automation checks and a ready-to-run skill for the three accelerator checks above.
Closing Thoughts
Like I said at the top, we all hear about AI making mistakes, and that it's on us to verify the output. That part stays true. But there are techniques and small adjustments to how you build the workflow that make the output more reliable and reduce the amount of verification it actually demands.
Tell me how it's going for you: are you getting reliable output, or are you still catching a lot of mistakes? Reply and let me know.
Next week is about something different: making AI outputs useful inside your thought process, not just reliable. You know that feeling when you're brainstorming with your favorite LLM, and it agrees with everything you suggest? That's where we're headed next week.
We Want Your Feedback!
This newsletter is for you, and we want to make it as valuable as possible. Please reply to this email with your questions, comments, or topics you'd like to see covered in future issues. Your input shapes our content!
Want to dive deeper into balanced AI adoption for your finance team? Or do you want to hire an AI-powered CFO? Book a consultation!
Did you find this newsletter helpful? Forward it to a colleague who might benefit!
Until next Tuesday, keep balancing!
Anna Tiomina
AI-Powered CFO
1
