Methodology
Benchmarking methodology
How Rise designs, runs, and scores Ficus benchmarks to measure AI productivity across professional domains.
Design principles
Real-world tasks
Every Ficus task is modeled on work that professionals actually perform. Tasks are authored by practicing domain experts whose daily work the tasks simulate.
Expert evaluation
AI outputs are evaluated by qualified professionals, not automated metrics alone. Evaluators are vetted for domain expertise and calibrated for consistency.
Economic significance
Tasks are selected for economic value — we measure whether AI can perform work that organizations currently pay professionals to do.
Contamination prevention
The full task set remains private to prevent models from being trained on benchmark data. An open subset is published for reproducibility.
Standardized scoring
All models are evaluated under identical conditions using the same scoring pipeline, ensuring fair and reproducible comparisons.
Evaluation process
Task Design
Domain experts design tasks that reflect real professional workflows, including multi-step reasoning, tool use, and cross-application work.
Model Evaluation
Frontier models complete tasks under standardized conditions. For agentic benchmarks, models interact with real tools in sandboxed environments.
Expert Scoring
Professional evaluators grade model outputs using rubrics designed to capture the quality standards of each domain.
Publication
Results are published on leaderboards with methodology documentation. New models are evaluated when they ship.
Get your model evaluated
Frontier labs and model developers can request evaluation through our partner form.
Request evaluation