Products
Infrastructure for the next generation of model training.
Copula Lab builds the benchmarks, expert workflow data, and RL environments that frontier labs use to push models into high-value, high-difficulty professional work.
Featured benchmarks
What we ship first.

GDPval
China GDP-weightedThe first productivity benchmark weighted by China’s official GDP structure.
GDPval measures how well frontier models produce real professional deliverables across Chinese economic tasks — not knowledge QA, but complete work products graded by domain experts. Task distribution is weighted by China’s official GDP industry structure and filtered by knowledge-work density, agent-oriented and mid-to-macro in granularity.
- Weighted by China’s official GDP industry structure
- Knowledge-work density filter, not raw GDP share
- Agent-oriented, long-horizon tasks (4h–1 day of expert work)
- Expert-authored rubrics + gold deliverables
- A leading indicator for AI’s labor-market impact in China

WebDev
Web · aesthetics & interactionCan agents build front-ends that actually look — and feel — good?
WebDev targets the real bottleneck: models can build front-ends, but not yet good ones. Each sample is a coding agent building a runnable, multi-file front-end from a natural-language brief, then graded by experts on aesthetics and complex interaction — taste, motion, context-fit, and AI-slop avoidance — not just whether it runs.
- Agentic: brief → runnable, multi-file front-end + screenshots
- Dimensional expert rubric (0–3), custom per-site anchors
- Aesthetics + interaction: layout, type, colour, motion, context-fit
- Anti-slop, ToB / ToC context-aware scoring
- Supports SFT · Preference / RL · Eval
What we offer
Four product lines, one supply chain.
Benchmarks
High-impact industry leaderboards that define the next wave of model deployment, build brand, and open procurement — scientific evaluation grounded in real economic tasks.
Expert RL Environments
High-value, long-horizon, semi-open task suites delivered as RL environments — task definitions, expert trajectories, automatic scoring functions, rubrics, and gold deliverables.
Expert Workflow Data
Domain data distilled from verified experts through AI-guided interviews and multi-expert QC, structured so models can learn, be evaluated, and be audited against how specialists reason.
Bad-Pattern Datasets
Datasets that map where frontier models break on long-chain professional tasks — fact citation, numerical consistency, evidence matching, implicit rules, and expert judgment.
Train on real expert work.
We work directly with model post-training, agent application training, and evaluation teams. Reach out to discuss benchmark access, RL environments, or expert data.
Talk to us