ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Mercor accounting study: frontier models ace every task while licensed CPAs average 37%

Mercor accounting study: frontier models ace every task while licensed CPAs average 37%

AI information • Admin • • 5 views

Mercor's accounting comparison study was published on October 2, 2026 by Mercor on its official blog: the team hired 12 junior accountants, all licensed CPAs with an average of five and a half years of experience, to face frontier AI models on the same month-end close tasks. The result is stark. The accountants averaged about 37%, while frontier models such as Claude Opus 5 scored perfectly on all 20 attempts, finishing each task in under 10 minutes against 30 to 180 minutes for the humans.

Not a quiz: digging through files for the right numbers

The tasks came from Mercor's earlier APEX-Accounting benchmark: four realistic month-end close scenarios. Participants had to dig through a company's working files to find the right numbers, do the math, and deliver a table of results. The scenarios hid easy-to-miss but realistic accounting requirements that compounded, so one overlooked figure could sink a score. The study was designed to measure how much AI lifts human performance, but the models scored perfectly unaided, leaving no uplift to measure, so Mercor reported the direct human-versus-model comparison instead.

The cost gap is wider than the speed gap

Per rubric criterion met, Claude Opus 5 cost about $0.21, against about $10.35 for the accountants at the US median wage, a gap of roughly 49 times. Eighteen months ago the best models could not reach the accountants' 37% average; OpenAI's o3 passed that line in spring 2025, GPT-5 reached about 69%, and frontier models now touch perfect scores. Some budget and open-weight models, such as Qwen3.5-122B, still sit below the accountant line.

A perfect score does not make accountants replaceable

Mercor draws its own boundaries clearly. The tasks test exactly what models are best at: searching files for detail and following instructions precisely. Client communication, asking the right questions, and accumulated business context, much of the real job, were not tested. The conditions also disadvantaged the humans: no colleagues to ask, no built-up context, and tasks the designers expected juniors to score only around 30% on anyway, since the benchmark was built to stump models and kept getting harder. What the study really shows is that structured, detail-heavy closing work is being taken over by models, while judgment, communication and review grow more valuable. The practical move for finance teams is to hand reconciliation, lookup and first drafts to models, and spend human time on checking and explaining.

Recommended Tools

More