Benchmark
Measure what models can do in real professional work.
Expert-designed benchmarks for evaluating frontier models on high-value, long-horizon tasks, scored against real deliverables and domain-specific rubrics.
Featured benchmarks
What we ship first.

GDPval
China GDP-weightedThe first productivity benchmark weighted by China’s official GDP structure.
GDPval measures how well frontier models produce real professional deliverables across Chinese economic tasks — not knowledge QA, but complete work products graded by domain experts. Task distribution is weighted by China’s official GDP industry structure and filtered by knowledge-work density, agent-oriented and mid-to-macro in granularity.
- Weighted by China’s official GDP industry structure
- Knowledge-work density filter, not raw GDP share
- Agent-oriented, long-horizon tasks (4h–1 day of expert work)
- Expert-authored rubrics + gold deliverables
- A leading indicator for AI’s labor-market impact in China

WebDev
Web · aesthetics & interactionCan agents build front-ends that actually look — and feel — good?
WebDev targets the real bottleneck: models can build front-ends, but not yet good ones. Each sample is a coding agent building a runnable, multi-file front-end from a natural-language brief, then graded by experts on aesthetics and complex interaction — taste, motion, context-fit, and AI-slop avoidance — not just whether it runs.
- Agentic: brief → runnable, multi-file front-end + screenshots
- Dimensional expert rubric (0–3), custom per-site anchors
- Aesthetics + interaction: layout, type, colour, motion, context-fit
- Anti-slop, ToB / ToC context-aware scoring
- Supports SFT · Preference / RL · Eval
How we evaluate
One method behind every benchmark.
Real tasks, real deliverables
Every task mirrors work a professional is paid to do. Models are scored on complete deliverables — documents, code, analyses — not multiple-choice answers.
Expert-authored rubrics
Domain experts define the scoring dimensions and per-task anchors, so a grade means what a specialist would mean by it.
Versioned methods and results
Each benchmark ships with its method, version history, and results, so scores stay comparable and reproducible as models improve.
Long-horizon, agent-oriented
Tasks span hours to days of expert work and are built for agents — multi-step, tool-using, judged end to end.
Train on real expert work.
We work directly with model post-training, agent application training, and evaluation teams. Reach out to discuss benchmark access, RL environments, or expert data.
Talk to us