copulaLAB

Products

Infrastructure for the next generation of model training.

Copula Lab builds the benchmarks, expert workflow data, and RL environments that frontier labs use to push models into high-value, high-difficulty professional work.

Featured benchmarks

What we ship first.

GDPval

GDPval

China GDP-weighted

The first productivity benchmark weighted by China’s official GDP structure.

GDPval measures how well frontier models produce real professional deliverables across Chinese economic tasks — not knowledge QA, but complete work products graded by domain experts. Task distribution is weighted by China’s official GDP industry structure and filtered by knowledge-work density, agent-oriented and mid-to-macro in granularity.

  • Weighted by China’s official GDP industry structure
  • Knowledge-work density filter, not raw GDP share
  • Agent-oriented, long-horizon tasks (4h–1 day of expert work)
  • Expert-authored rubrics + gold deliverables
  • A leading indicator for AI’s labor-market impact in China
Talk to us
WebDev

WebDev

Web · aesthetics & interaction

Can agents build front-ends that actually look — and feel — good?

WebDev targets the real bottleneck: models can build front-ends, but not yet good ones. Each sample is a coding agent building a runnable, multi-file front-end from a natural-language brief, then graded by experts on aesthetics and complex interaction — taste, motion, context-fit, and AI-slop avoidance — not just whether it runs.

  • Agentic: brief → runnable, multi-file front-end + screenshots
  • Dimensional expert rubric (0–3), custom per-site anchors
  • Aesthetics + interaction: layout, type, colour, motion, context-fit
  • Anti-slop, ToB / ToC context-aware scoring
  • Supports SFT · Preference / RL · Eval
Talk to us

What we offer

Four product lines, one supply chain.

Benchmarks

High-impact industry leaderboards that define the next wave of model deployment, build brand, and open procurement — scientific evaluation grounded in real economic tasks.

Expert RL Environments

High-value, long-horizon, semi-open task suites delivered as RL environments — task definitions, expert trajectories, automatic scoring functions, rubrics, and gold deliverables.

Expert Workflow Data

Domain data distilled from verified experts through AI-guided interviews and multi-expert QC, structured so models can learn, be evaluated, and be audited against how specialists reason.

Bad-Pattern Datasets

Datasets that map where frontier models break on long-chain professional tasks — fact citation, numerical consistency, evidence matching, implicit rules, and expert judgment.

Train on real expert work.

We work directly with model post-training, agent application training, and evaluation teams. Reach out to discuss benchmark access, RL environments, or expert data.

Talk to us