The AI Productivity Index for Finance

The AI Productivity Index for Finance (Ficus-Finance) measures whether frontier AI agents can execute long-horizon, cross-application tasks in professional accounting and financial services.

The Ficus-Finance leaderboard

We created Ficus-Finance to evaluate agents on the real day-to-day work of finance professionals: auditors, tax specialists, and financial analysts. The tasks require agents to navigate accounting software, spreadsheets, and PDFs while maintaining professional standards.

Ficus-Finance was built with practicing professionals who created realistic scenarios spanning audit preparation, tax reconciliation, and financial reporting. Each task mirrors the complexity and multi-step nature of actual professional work.

The benchmark evaluates both accuracy and completeness, ensuring that AI models can handle the precision required in professional accounting.

Model

Score

Opus 5.5Max

61.8% ±5.2%

Fable 5.1Max

61% ±5.3%

Opus 5.5Medium

59.9% ±5.4%

Gemini 4 ArgonHigh

57.3% ±5.6%

Sonnet 5.5Max

55.8% ±5.7%

0%
20%
40%
60%
80%
100%

Domains evaluated in Ficus-Finance

Audit Preparation

Tests AI on preparing audit workpapers, reviewing financial statements, identifying discrepancies, and ensuring compliance with accounting standards.

Tax Reconciliation

Evaluates AI ability to reconcile tax obligations, review deductions, and prepare accurate tax filings across multiple jurisdictions.

Financial Reporting

Assesses AI on generating financial reports, analyzing variance, and producing management-ready summaries from raw accounting data.

Frequently asked questions

What is the Ficus-Finance benchmark?

Ficus-Finance is a benchmark that measures how effectively AI models perform professional accounting and financial tasks. It evaluates long-horizon, multi-step workflows across audit preparation, tax reconciliation, and financial reporting.

Who creates the Ficus-Finance tasks?

Practicing finance professionals create the tasks, including CPAs, auditors, and financial analysts with experience at leading accounting and advisory firms.

How are models evaluated?

Ficus-Finance evaluates the quality of completed financial work. Model outputs are graded using expert-authored rubrics, measuring accuracy, completeness, and compliance with accounting standards.

Can I reproduce results?

Yes, on the open subset. The eval harness and sample tasks are available for independent verification. The full task set stays private to prevent training contamination.

← Back to Ficus Benchmarks