Skip to main content
Accounting

AI beats out CPAs at month-end close tasks

Opus 5.5 got 100% of tasks correct on a benchmark, while human accountants only averaged 37%.

• 3 min read

TOPICS: Accounting / Emerging Trends / Automation Impact

Forget rogue agents. Accountants might want to be looking over their shoulders for the plain-vanilla LLMs.

Claude’s Opus 5 outperformed CPAs on a benchmark test designed to mimic a real month-end close situation, according to AI training firm Mercor, which developed the test. The AI got 100% of the items on the rubric correct, while human accountants got only an average of 37%.

“Models today already seem more meticulous than junior accountants,” Mercor’s researchers wrote in a report. “That’s extremely valuable in a field where small errors can be extraordinarily costly.”

And Opus 5 was also faster, completing the tasks in under 15 minutes, while the humans took from 30 minutes to 3 hours.

The AI was cheaper as well. Mercor compared the token costs of running the LLM to the BLS median salary for accountants, and found the Opus 5 could complete an item on the rubric correctly for $0.21, versus $10.35 for human accountants.

They’ve caught up. The human accountants weren’t exactly slouches. The 12 flesh-and-blood participants were all US-based practicing accountants with active CPA licenses with an average of 5.4 years’ experience. Six of them had worked for Big Four firms.

The tasks they were asked to perform were fairly complex. Developed by accounting and finance professionals, the benchmark asked participants to look through fictional organizations’ working files to find the right numbers, “figure out which calculation approach the company’s rules called for, do the math, and then deliver a completed summary table of the results.” The test included such realistic and “not rare” problems as “something coded to the wrong account, a bill that lands after the month ends, [or] a commission that never made into the books,” one of the Mercor test developers noted.

News built for finance pros

CFO Brew helps finance pros navigate their roles with insights into risk management, compliance, and strategy through our newsletter, virtual events, and digital guides.

By subscribing, you accept our Terms & Privacy Policy.

What’s perhaps most unsettling is the fact that, 18 months ago, even the best AI models weren’t able to beat human accountants. “These data do suggest substantial changes, and productivity gains, in accounting in the coming years,” Mercor researcher Aden Barton said in a blog post. “That would be true even if model progress froze today.”

A few caveats: The entries were scored by an LLM judge—though when humans hand-graded some of the tasks, they agreed with the bot 98% of the time, the Mercor report said.

And the nature of the test may have made it more favorable to AI than to humans: “Model outperformance came not from the task difficulty per se but from its interaction with the setting and the scoring,” the researchers stated. The CPAs weren’t able to ask questions, as they would in the real world, and the tasks “tested specific concepts that may not come up regularly depending on an accountant’s specialization.”

Five out of 12 of the participants “said the tasks were not realistic.”

The authors’ conclusion? Advances in AI may mean that “accountants’ work will likely shift toward the parts of the job where speed and meticulousness matter less.” Good thing accountants aren’t known for their meticulousness.

About the author

Courtney Vien

Courtney Vien is a senior reporter for CFO Brew. She formerly served as editor in chief of the Journal of Accountancy.

CFO Brew helps finance pros navigate their roles with insights into risk management, compliance, and strategy through our newsletter, virtual events, and digital guides.

By subscribing, you accept our Terms & Privacy Policy.