Methodology

Benchmarking methodology

How Rise designs, runs, and scores Ficus benchmarks to measure AI productivity across professional domains.

Design principles

01

Real-world tasks

Every Ficus task is modeled on work that professionals actually perform. Tasks are authored by practicing domain experts whose daily work the tasks simulate.

02

Expert evaluation

AI outputs are evaluated by qualified professionals, not automated metrics alone. Evaluators are vetted for domain expertise and calibrated for consistency.

03

Economic significance

Tasks are selected for economic value — we measure whether AI can perform work that organizations currently pay professionals to do.

04

Contamination prevention

The full task set remains private to prevent models from being trained on benchmark data. An open subset is published for reproducibility.

05

Standardized scoring

All models are evaluated under identical conditions using the same scoring pipeline, ensuring fair and reproducible comparisons.

Evaluation process

1

Task Design

Domain experts design tasks that reflect real professional workflows, including multi-step reasoning, tool use, and cross-application work.

2

Model Evaluation

Frontier models complete tasks under standardized conditions. For agentic benchmarks, models interact with real tools in sandboxed environments.

3

Expert Scoring

Professional evaluators grade model outputs using rubrics designed to capture the quality standards of each domain.

4

Publication

Results are published on leaderboards with methodology documentation. New models are evaluated when they ship.

Get your model evaluated

Frontier labs and model developers can request evaluation through our partner form.

Request evaluation