Ficus Benchmarks
The Ficus family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.
Our benchmarks
Each benchmark tests a different dimension of professional capability.
Ficus-Workflows 1.1
Long-horizon, cross-application tasks in professional services
Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.
Ficus-Finance
Long-horizon, cross-application tasks in professional accounting
Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.
Ficus-Code
Real-world software engineering across integration and observability
Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds.
Off-the-shelf data
License-ready datasets, expert-written and graded.
FAQ
What are Rise's Ficus benchmarks and how do they work?
The Rise AI Productivity Index (Ficus) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. It provides data-driven, real-world productivity metrics across high-value sectors, like software engineering, corporate law, investment banking, accounting, and management consulting. The suite includes benchmarks such as Ficus-Workflows, which evaluates long-horizon, multi-step agent workflows; Ficus-Code, focused on software engineering; and Ficus-Finance, focused on agentic accounting tasks.
Which model scores highest on Ficus?
Rankings can change whenever a frontier model is released. New models are evaluated on Ficus-Workflows, Ficus-Code, and Ficus-Finance when they ship.
Who writes the Ficus tasks?
Practicing professionals whose daily work the tasks simulate. Task authors are vetted domain experts across law, consulting, banking, accounting, medicine, and software engineering.
Can I reproduce Ficus results myself?
Yes, on the open subset. The eval harness and sample tasks are published so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.
How do I get my model evaluated on Ficus?
Frontier labs and model developers can request evaluation through our partner form.
Learn more →Can I license Ficus data for training?
Yes. Rise licenses off-the-shelf datasets built by the same expert network, with samples available the same day.
Learn more →Ficus Newsletter
The latest on frontier AI performance, straight to your inbox.
New benchmarks, leaderboard shifts, and research from the Ficus team.
By subscribing you agree to receive updates from Rise. Unsubscribe anytime.