Post-Training Research Scientist

Two Sigma Investments
New York, NY
Category Research
Job Description
Role Overview

We are hiring a Post-Training Research Scientist to build RLHF, DPO, and reward modeling capabilities from the ground up. This is a greenfield role: you will define the infrastructure, research agenda, and evaluation frameworks for aligning LLMs to sophisticated, multi-step workflows in a domain where the reward signal is fundamentally different from existing research on human preference or deterministic task completion.

What You Will Do

Lead post-training efforts for LLMs applied to financial time series and quantitative reasoning, design and execute RLHF, DPO, and related alignment methods at scale, build infrastructure for preference data collection, reward modeling, and policy optimization on financial datasets, drive research agenda connecting post-training methods to quantitative finance applications, collaborate with quant researchers to define task distributions and evaluation frameworks, unblock production systems dependent on post-training capabilities.

Why It Might Be a Fit

You should possess a BS or equivalent work experience in Science, Technology, Engineering or Math, minimum 1 year of experience required; 1-10 years of experience preferred (ideally 1-5 years) at a frontier AI lab, shipped post-training systems in production, deep understanding of distributed training infrastructure, track record managing large-scale compute, publications or demonstrated expertise in alignment, preference learning, or reward modeling, hands-on implementation skills: PyTorch/JAX, distributed frameworks (DeepSpeed, FSDP, etc.).

Requirements

  • BS or equivalent work experience in Science, Technology, Engineering or Math
  • Minimum 1 year of experience required; 1-10 years of experience preferred (ideally 1-5 years) at a frontier AI lab
  • Shipped post-training systems in production: RLHF, DPO, RLAIF, or related methods
  • Deep understanding of distributed training infrastructure: multi-node GPU clusters, training stability, checkpointing
  • Track record managing large-scale compute: budgeting, experiment design, ablations
  • Publications or demonstrated expertise in alignment, preference learning, or reward modeling
  • Hands-on implementation skills: PyTorch/JAX, distributed frameworks (DeepSpeed, FSDP, etc.)

Benefits

  • Fully paid medical and dental insurance premiums for employees and dependents
  • Competitive 401k match
  • Employer-paid life & disability insurance
  • Onsite gyms with laundry service
  • Wellness activities
  • Casual dress
  • Snacks
  • Game rooms
  • Tuition reimbursement
  • Conference and training sponsorship
  • Generous vacation and unlimited sick days
  • Competitive paid caregiver leaves
  • Flexible in-office days with budget for home office setup
]]>