The AI failures that reached finance over the past two years were failures of review. Somebody accepted an output without checking it, and the fix was a person doing their job properly. July produced a different kind, where an agent was given access, used it on its own, and moved through a company’s live systems before anyone noticed.
Two surveys published nine days apart put numbers on how close that sits to us, and they read badly together. Many finance teams have already given agents the ability to create records and approve transactions, yet most cannot verify afterward what the agent actually did. Meanwhile, the pressure to deploy is arriving from people who read the same headlines you do.
Four items this month, with my take on each. For paid subscribers, I have written the thing I get asked about most: where I would actually put an agent across six finance processes, exactly where I would stop, and what you can show for it in six weeks.
Cohort 2 of Claude in Action wraps up today, and Cohort 3 starts September 4.
Over four sessions, we build toward automating your workflow end-to-end. What participants have built so far:
Intercompany AR reconciliation project
A skill that turns a raw ERP export into a dashboard
A board reporting pack that runs from a single command
An ASC 842 lease worksheet speeding up the close
You get there by building on your own data while I check your work, so seats are capped, and the group stays small. I am holding the rate at the early promotional level.
Sign up here.
I can also run Claude in Action inside your team, on your own workflows and data examples, with everything we build staying with you. Reply to this email, and we can work out whether it fits.
What Happened in July
An AI agent broke into a company’s live systems by itself
In mid-July, an AI model that OpenAI was running through a security test went past the limits it had been given, reached the open internet, and got inside Hugging Face’s live systems, where it collected access credentials and used them to move from one internal system to another. Hugging Face disclosed it on July 21 and only reconstructed the sequence afterward.
This belongs in a finance newsletter for a specific reason. Through 2025, the AI failures that reached us were failures of review, where somebody accepted an output without checking it, and the fix was a person doing their job properly. This one is different in a way that matters to anyone who designs controls. The agent had access; it acted on that access without being told to, and nobody saw it while it was happening. Segregation of duties, approval limits and review controls all assume a person takes an action and a second person checks it; an agent holding system access is a third party those controls were never written for.
Two surveys published in the same week read badly against each other. Pathlock found that 38% of organizations let AI agents create or modify business records and 28% let them approve transactions, while 52% cannot verify what their agents actually did and 79% have nobody who owns AI governance. Avalara found that 92% of finance leaders are under career pressure to show ROI on agents, only 7% put governance ahead of deployment speed, and 76% have no one in-house who understands how their own agents work.
So what: I adopt AI early, and I push my clients to adopt, so it costs me something to write this, but agents are not ready for critical workflows; the verification has not caught up with the access. Give them work where a mistake is cheap and a check takes minutes, and keep them away from anything that posts, pays, or approves.
Claude can now learn a task by watching you work
Two Cowork changes landed in July. On July 7 Cowork moved to web and mobile, with sessions running on Anthropic’s servers instead of your machine, so a scheduled task can run with no device powered on. On July 21 Record a Skill arrived in Cowork on the Mac desktop app, where you narrate a screen recording of a task and Claude turns it into a skill it can run again.
The catch is what a recording cannot pick up. It captures the steps you took, not the moments you would have stopped, and the check you run without noticing you are running it is exactly the control that never makes it into the skill.
So what: you can now build a working skill without writing a prompt, which removes the main reason finance people never build one. Record something you do the same way every month, then go back through it and add the checks you make automatically, because those are the parts the recording misses.
Token prices fell while seat prices rose
Anthropic released Opus 5 on July 24 and OpenAI released GPT-5.6 on July 9 in three versions of different strength and price. Then prices moved, and not in one direction.
Change | Date | Old | New |
|---|---|---|---|
GPT-5.6 Luna, the cheapest tier | Jul 30 | $1.00 / $6.00 | $0.20 / $1.20 |
GPT-5.6 Terra, the middle tier | Jul 30 | $2.50 / $15.00 | $2.00 / $12.00 |
Gemini 3.6 Flash, output | Jul 21 | $9.00 | $7.50 |
Claude Sonnet 5, intro to Aug 31 | Jul 1 | n/a | $2.00 / $10.00 |
Claude Fable 5, now metered | Jul 7 | included | $10.00 / $50.00 |
Microsoft 365 E3 | Jul 1 | $36 per seat | $39 per seat |
Microsoft 365 E5 | Jul 1 | $57 per seat | $60 per seat |
Model rows are the price per million tokens in and out, which is roughly 750,000 words each way; Microsoft rows are per user per month. The strongest OpenAI tier got no cut at all. Microsoft did more than raise the seat price, since storage is now billed as you use it, so cost that used to sit inside the subscription now shows up on the invoice, the new E7 bundle is $99 a seat, and Copilot Cowork needs a licence plus usage credits on top of it.
So what: your AI cost is now a fixed seat charge plus a variable usage charge, so budget it as two lines rather than one. Check which model your team runs by default, because after July 30 putting everyday work through the strongest model can cost ten times what a middle one does, and Sonnet 5’s introductory rate ends August 31.
What EU enforcement could do to your automation
August 2, enforcement powers under the EU AI Act switched on for companies, and the piece that reaches you directly is Article 50 transparency, so anyone interacting with your AI has to be told it is AI, and AI-generated content has to be labeled. The high-risk rules most finance teams prepared for, credit scoring among them, moved to December 2027.
The question I keep getting is a different one: what happens to a workflow built on Claude or ChatGPT if the Act is enforced against the provider? Article 93 lets the Commission require a provider to restrict availability, withdraw, or recall a model where there is serious systemic risk. A full withdrawal is unlikely, and the likelier version is duller, since a study of 375 models found 11% of advanced releases delayed or blocked outright in the EU compared with the US, with Claude 3 Opus arriving 71 days late.
So what: this is vendor concentration risk in new clothing. Check whether your process can run on a different model, and have a backup plan for your AI automations.
The subscriber section is usually about “What to do”. For this month’s news, the main question is: what do you actually do when agents are going past their limits on one side, and the pressure to implement them is coming at you from the other?
My answer is to choose the work rather than argue with the pressure. In the subscriber section: where I would put an agent across six finance processes, accounts payable through reporting, the exact point I would stop in each one, and what you can show for it in six weeks.
Closing Thoughts
If you take one thing from this month, make it the distinction between an AI that produces something you check and an agent that does something you find out about later. Everything else follows from which of those you are buying.
I read every reply, so if you are further along on agents than this edition suggests you should be, tell me what is working.
We Want Your Feedback!
This newsletter is for you, and we want to make it as valuable as possible. Please reply to this email with your questions, comments, or topics you'd like to see covered in future issues. Your input shapes our content!
Want to dive deeper into balanced AI adoption for your finance team? Or do you want to hire an AI-powered CFO? Book a consultation!
Did you find this newsletter helpful? Forward it to a colleague who might benefit!
Until next Tuesday, keep balancing!
Anna Tiomina
AI-Powered CFO
1
