← All roles
Research · Data
Research Scientist
This is not a traditional model-tuning or applied-ML role. We need someone who understands frontier capability iteration — post-training, eval, agent training, and data strategy — and can also analyze how models fail in real business scenarios. Your core output is research judgment, evaluation methods, data strategy, and experimental conclusions. You’ll independently own a research or data direction, turning fuzzy problems into clear, deliverable benchmarks, data products, or training-feedback plans.
What you’ll do
- 01Track leading model teams’ needs and movements across post-training, agent training, eval, RL environments, and synthetic data.
- 02Build an iteration loop from training feedback to data strategy, grounded in model evals, customer feedback, and error analysis.
- 03Independently own a research or data direction — problem definition, literature review, experiment design, data construction, evaluation methods, and retrospectives.
- 04Systematically analyze model failure modes in real tasks: understanding, reasoning, tool use, multi-turn interaction, visual recognition, professional judgment, task planning, and result reliability.
- 05Design benchmarks, evals, rubrics, quality standards, and training-data plans that genuinely serve capability gains.
- 06Work with domain experts, algorithm, data-engineering, and delivery teams to decompose expert judgment and real workflows into learnable, evaluable task units.
What we look for
- 01Master’s or above in AI, CS, statistics, math, or a related field.
- 02Experience with LLMs, NLP, multimodal, agents, model eval, post-training, data construction, or training-data analysis.
- 03Understanding of SFT, RLHF / RLAIF, preference data, reward models, eval, agent training, and synthetic data.
- 04Solid research training — paper reading, experiment design, data analysis, error attribution, and technical writing.
- 05Strong coding ability — Python and common ML / data tools; able to build model calls, data processing, eval scripts, and experiment analysis.
- 06Strong ownership — able to break fuzzy problems into clear goals, experimental paths, and deliverables.
Nice to have
- Conference papers, high-quality arXiv / workshop work, open-source projects, or research internships.
- Experience at a top AI lab, big-tech model team, research institution, or strong startup.
- Experience with post-training, benchmarks, eval, agent / multimodal / code evaluation, or high-quality dataset construction.
- Your own views on capability boundaries, synthetic data, agent training, RL environments, or expert data.
Apply for this role
Send your résumé and a short note on why this role.