The ongoing discussion regarding artificial intelligence in corporate finance has received a concrete empirical baseline. On October 1, 2026, talent and evaluation platform Mercor released the findings of a study conducted by author Aden Barton, built entirely around the APEX-Accounting Benchmark. The test directly evaluated human certified public accountants against a state of the art frontier model on realistic, document-heavy accounting challenges. The resulting performance metrics offer a rigorous look at how automated workflows might reshape audit and accounting practices.
The study evaluated twelve licensed US certified public accountants possessing an average of 5.4 years of professional industry experience. These professionals were tasked with resolving four realistic month-end close scenarios, which encompassed intricate lease schedules alongside detailed revenue and commission reconciliations. These scenarios represent the core operational duties performed routinely by finance teams in corporate environments. Industry practitioners have traditionally assumed that such nuanced, cross-document reconciliation tasks require seasoned human judgment to avoid critical balance sheet discrepancies.
However, the performance of the human accountants revealed surprising vulnerabilities in manual execution. Across the four scenarios, the participating CPAs achieved an average accuracy rate of only 37 percent. Completing the review tasks demanded substantial time, with individual sessions ranging from 30 to 180 minutes per assignment. This high expenditure of billable hours drove the operational cost to an average of approximately 10.35 dollars per evaluated criterion, illustrating the financial weight of conventional manual workflows.
In contrast, the frontier model Claude Opus 5 delivered an entirely autonomous and flawless performance across the identical test battery. The model completed every assignment with a 100 percent success rate, concluding the entire verification process in under ten minutes. The computational expense for the system amounted to merely 0.21 dollars per criterion, representing a steep drop in direct resource utilization. The stark difference underlines both the speed and the potential cost containment offered by advanced agentic architectures.
These findings deliver an early empirical confirmation that agentic workflows can outperform experienced human specialists in complex document reconciliations rather than just basic repetitive data entry. While legacy automation in finance has targeted low-complexity bookkeeping entries, the APEX benchmark shows that autonomous systems can substantially undercut human error rates during demanding monthly close procedures. The findings suggest that complex reconciliation processes, previously considered safe from total automation, are highly susceptible to disruption.
For corporate accounting divisions and audit institutions, these findings indicate an impending structural transformation in operating models. As autonomous models prove capable of resolving intricate leasing schedules and commission alignments with zero recorded errors, the daily responsibilities of accounting professionals will inevitably shift toward supervisory oversight and model governance. The focus will migrate away from tedious spreadsheet verifications toward verifying agent outputs and interpreting broader balance sheet dynamics. The benchmark published by Mercor provides a clear signal that manual document checking is nearing an evolutionary turning point.

