The AI Productivity Index for Finance
The AI Productivity Index for Finance (Ficus-Finance) measures whether frontier AI agents can execute long-horizon, cross-application tasks in professional accounting and financial services.
The Ficus-Finance leaderboard
We created Ficus-Finance to evaluate agents on the real day-to-day work of finance professionals: auditors, tax specialists, and financial analysts. The tasks require agents to navigate accounting software, spreadsheets, and PDFs while maintaining professional standards.
Ficus-Finance was built with practicing professionals who created realistic scenarios spanning audit preparation, tax reconciliation, and financial reporting. Each task mirrors the complexity and multi-step nature of actual professional work.
The benchmark evaluates both accuracy and completeness, ensuring that AI models can handle the precision required in professional accounting.
Model
Score
Opus 5.5Max
61.8% ±5.2%
Fable 5.1Max
61% ±5.3%
Opus 5.5Medium
59.9% ±5.4%
Gemini 4 ArgonHigh
57.3% ±5.6%
Sonnet 5.5Max
55.8% ±5.7%
Audit Preparation
Tests AI on preparing audit workpapers, reviewing financial statements, identifying discrepancies, and ensuring compliance with accounting standards.
Tax Reconciliation
Evaluates AI ability to reconcile tax obligations, review deductions, and prepare accurate tax filings across multiple jurisdictions.
Financial Reporting
Assesses AI on generating financial reports, analyzing variance, and producing management-ready summaries from raw accounting data.
Frequently asked questions
What is the Ficus-Finance benchmark?
Ficus-Finance is a benchmark that measures how effectively AI models perform professional accounting and financial tasks. It evaluates long-horizon, multi-step workflows across audit preparation, tax reconciliation, and financial reporting.
Who creates the Ficus-Finance tasks?
Practicing finance professionals create the tasks, including CPAs, auditors, and financial analysts with experience at leading accounting and advisory firms.
How are models evaluated?
Ficus-Finance evaluates the quality of completed financial work. Model outputs are graded using expert-authored rubrics, measuring accuracy, completeness, and compliance with accounting standards.
Can I reproduce results?
Yes, on the open subset. The eval harness and sample tasks are available for independent verification. The full task set stays private to prevent training contamination.