The AI Productivity Index for Agents
The AI Productivity Index for Agents (Ficus-Workflows) measures whether frontier AI agents can execute long-horizon, cross-application tasks across three jobs in professional services.
The Ficus-Workflows leaderboard
We created Ficus-Workflows to evaluate agents on the real day-to-day work of professionals: investment banking analysts, management consultants, and corporate lawyers. The tasks require agents to reason, demonstrate advanced knowledge, use multiple applications, and plan over long horizons.
Ficus-Workflows was built in three steps. First, industry professionals created a data-rich world, based on a unique project scenario. Second, they created realistic, challenging tasks using the files from within the world. Third, we gave agents access so they could execute the tasks.
The entire Ficus-Workflows dataset is available open-source, along with our infrastructure for executing and evaluating agent trajectories.
Model
Score
Gemini 4 ArgonHigh
82.2% ±4.4%
Sonnet 5.5Max
75.5% ±4.7%
Opus 5.5Max
73.5% ±4.9%
Fable 5.1Max
68.6% ±4.9%
Gemini 3.7 FlashHigh
67.8% ±5.1%
Corporate Lawyer
Long-horizon legal tasks built by practicing attorneys. Evaluates AI on legal research, contract drafting, due diligence review, and regulatory analysis.
Management Consultant
Multi-step research and strategic analysis tasks, built by consultants. Tests AI on market analysis, competitive intelligence, and strategic recommendations.
Investment Banking Analyst
Multi-step investment banking workflows evaluating AI on financial modeling, valuation, deal structuring, and pitch preparation.
Frequently asked questions
What is the Ficus-Workflows benchmark and how does it work?
Ficus-Workflows is part of the Ficus family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. It evaluates long-horizon, multi-step workflows across high-value sectors like corporate law, investment banking, and management consulting.
Which AI model scores highest on Ficus-Workflows?
Rankings change whenever a frontier model is released. New models are evaluated on all Ficus benchmarks when they ship.
How are AI models evaluated on Ficus-Workflows?
Ficus-Workflows evaluates the quality of completed work. Model outputs are graded using expert-authored rubrics with an LM judge.
What do Mean Score and Pass@1 measure?
Mean Score measures the average percentage of criteria that are passed for all tasks in the benchmark. Pass@1 measures the percentage of tasks that an agent successfully completes on a single attempt. Together, they capture overall capability and end-to-end task reliability.
Can I reproduce Ficus-Workflows results myself?
Yes, on the open subset. The eval harness and sample tasks are available so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.