Ficus Benchmarks

The Ficus family of benchmarks assesses whether frontier AI models can perform economically valuable tasks across professional services, medicine, and software engineering.

Our benchmarks

Each benchmark tests a different dimension of professional capability.

Ficus-Workflows 1.1

Long-horizon, cross-application tasks in professional services

Tests whether AI agents can complete multi-hour professional tasks across investment banking, corporate law, and management consulting, using real tools in Google Suite.

View Leaderboard
FicusView leaderboard →

Ficus-Finance

Long-horizon, cross-application tasks in professional accounting

Measuring AI agents ability to complete professional accounting tasks across tools like accounting software, spreadsheets, and PDFs.

View Leaderboard
FicusView leaderboard →

Ficus-Code

Real-world software engineering across integration and observability

Measures AI performance on real-world software engineering tasks, from bug fixes to feature builds.

View Leaderboard
FicusView leaderboard →
New

Off-the-shelf data

License-ready datasets, expert-written and graded.

Learn more

FAQ

What are Rise's Ficus benchmarks and how do they work?

The Rise AI Productivity Index (Ficus) is a family of benchmarks that measure how effectively AI models and agents perform economically valuable tasks. It provides data-driven, real-world productivity metrics across high-value sectors, like software engineering, corporate law, investment banking, accounting, and management consulting. The suite includes benchmarks such as Ficus-Workflows, which evaluates long-horizon, multi-step agent workflows; Ficus-Code, focused on software engineering; and Ficus-Finance, focused on agentic accounting tasks.

Which model scores highest on Ficus?

Rankings can change whenever a frontier model is released. New models are evaluated on Ficus-Workflows, Ficus-Code, and Ficus-Finance when they ship.

Who writes the Ficus tasks?

Practicing professionals whose daily work the tasks simulate. Task authors are vetted domain experts across law, consulting, banking, accounting, medicine, and software engineering.

Can I reproduce Ficus results myself?

Yes, on the open subset. The eval harness and sample tasks are published so you can run the same scoring pipeline against your own model. The full task set stays private so that models can't be trained on it.

How do I get my model evaluated on Ficus?

Frontier labs and model developers can request evaluation through our partner form.

Learn more →
Can I license Ficus data for training?

Yes. Rise licenses off-the-shelf datasets built by the same expert network, with samples available the same day.

Learn more →

Ficus Newsletter

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the Ficus team.

By subscribing you agree to receive updates from Rise. Unsubscribe anytime.