copulaLAB

Benchmark

Measure what models can do in real professional work.

Expert-designed benchmarks for evaluating frontier models on high-value, long-horizon tasks, scored against real deliverables and domain-specific rubrics.

Featured benchmarks

What we ship first.

GDPval

GDPval

China GDP-weighted

The first productivity benchmark weighted by China’s official GDP structure.

GDPval measures how well frontier models produce real professional deliverables across Chinese economic tasks — not knowledge QA, but complete work products graded by domain experts. Task distribution is weighted by China’s official GDP industry structure and filtered by knowledge-work density, agent-oriented and mid-to-macro in granularity.

  • Weighted by China’s official GDP industry structure
  • Knowledge-work density filter, not raw GDP share
  • Agent-oriented, long-horizon tasks (4h–1 day of expert work)
  • Expert-authored rubrics + gold deliverables
  • A leading indicator for AI’s labor-market impact in China
Talk to us
WebDev

WebDev

Web · aesthetics & interaction

Can agents build front-ends that actually look — and feel — good?

WebDev targets the real bottleneck: models can build front-ends, but not yet good ones. Each sample is a coding agent building a runnable, multi-file front-end from a natural-language brief, then graded by experts on aesthetics and complex interaction — taste, motion, context-fit, and AI-slop avoidance — not just whether it runs.

  • Agentic: brief → runnable, multi-file front-end + screenshots
  • Dimensional expert rubric (0–3), custom per-site anchors
  • Aesthetics + interaction: layout, type, colour, motion, context-fit
  • Anti-slop, ToB / ToC context-aware scoring
  • Supports SFT · Preference / RL · Eval
Talk to us

How we evaluate

One method behind every benchmark.

Real tasks, real deliverables

Every task mirrors work a professional is paid to do. Models are scored on complete deliverables — documents, code, analyses — not multiple-choice answers.

Expert-authored rubrics

Domain experts define the scoring dimensions and per-task anchors, so a grade means what a specialist would mean by it.

Versioned methods and results

Each benchmark ships with its method, version history, and results, so scores stay comparable and reproducible as models improve.

Long-horizon, agent-oriented

Tasks span hours to days of expert work and are built for agents — multi-step, tool-using, judged end to end.

Train on real expert work.

We work directly with model post-training, agent application training, and evaluation teams. Reach out to discuss benchmark access, RL environments, or expert data.

Talk to us