The AI Productivity Index for Agents

The AI Productivity Index for Agents (Ficus-Workflows) measures whether frontier AI agents can execute long-horizon, cross-application tasks across three jobs in professional services.

The Ficus-Workflows leaderboard

We created Ficus-Workflows to evaluate agents on the real day-to-day work of professionals: investment banking analysts, management consultants, and corporate lawyers. The tasks require agents to reason, demonstrate advanced knowledge, use multiple applications, and plan over long horizons.

Ficus-Workflows was built in three steps. First, industry professionals created a data-rich world, based on a unique project scenario. Second, they created realistic, challenging tasks using the files from within the world. Third, we gave agents access so they could execute the tasks.

The entire Ficus-Workflows dataset is available open-source, along with our infrastructure for executing and evaluating agent trajectories.

Model

Score

Gemini 4 ArgonHigh

82.2% ±4.4%

Sonnet 5.5Max

75.5% ±4.7%

Opus 5.5Max

73.5% ±4.9%

Fable 5.1Max

68.6% ±4.9%

Gemini 3.7 FlashHigh

67.8% ±5.1%

0%
20%
40%
60%
80%
100%

Domains evaluated in Ficus-Workflows

Corporate Lawyer

Long-horizon legal tasks built by practicing attorneys. Evaluates AI on legal research, contract drafting, due diligence review, and regulatory analysis.

Management Consultant

Multi-step research and strategic analysis tasks, built by consultants. Tests AI on market analysis, competitive intelligence, and strategic recommendations.

Investment Banking Analyst

Multi-step investment banking workflows evaluating AI on financial modeling, valuation, deal structuring, and pitch preparation.

Frequently asked questions

What is the Ficus-Workflows benchmark and how does it work?

Ficus-Workflows is part of the Ficus family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. It evaluates long-horizon, multi-step workflows across high-value sectors like corporate law, investment banking, and management consulting.

Which AI model scores highest on Ficus-Workflows?

Rankings change whenever a frontier model is released. New models are evaluated on all Ficus benchmarks when they ship.

How are AI models evaluated on Ficus-Workflows?

Ficus-Workflows evaluates the quality of completed work. Model outputs are graded using expert-authored rubrics with an LM judge.

What do Mean Score and Pass@1 measure?

Mean Score measures the average percentage of criteria that are passed for all tasks in the benchmark. Pass@1 measures the percentage of tasks that an agent successfully completes on a single attempt. Together, they capture overall capability and end-to-end task reliability.

Can I reproduce Ficus-Workflows results myself?

Yes, on the open subset. The eval harness and sample tasks are available so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.

← Back to Ficus Benchmarks