Research Engineer, Post-Training
About the role
You will own the methods that turn production feedback into model improvements: supervised fine-tuning on curated trajectories, preference and reward modeling from human corrections, and offline and online RL against task environments built from real customer workflows.
This is an applied role. Experiments are judged by whether they move held-out task success for a real customer, and you will be expected to understand the data as deeply as the algorithm.
What you'll do
- Design, run, and analyze post-training experiments on 7B to 70B+ parameter models
- Build reward signals and environments from customer tool-use data
- Diagnose regressions: when a new checkpoint is worse at something, find out why
- Write clear internal reports that let the rest of the team make decisions
- Contribute to shared training infrastructure and evaluation tooling
Requirements
- Deep familiarity with modern LLM post-training (SFT, DPO/RLHF-style methods, GRPO or similar)
- Strong PyTorch skills and experience with distributed training
- A track record of getting experiments to a conclusion, not just to a plot
- Comfort working with imperfect, real-world data
Nice to have
- Publications or open-source work in RL, agents, or evaluation
- Experience with multi-turn tool-use environments
Compensation & benefits
- Base salary: $200,000 to $320,000
- Meaningful early-stage equity
- Medical, dental, and vision coverage
- 401(k)
- Relocation support to San Francisco
- Daily lunch in the office and a learning stipend
About Clad Labs
Clad Labs is an applied AI lab building systems to train, evaluate, and continuously improve production agents. We turn the data a company already generates, real requests, tool calls, and human corrections, into training signal, and we measure every change against held-out tasks before it ships.
We are a small team working in person in San Francisco. We care about doing careful work on hard problems, and about being honest with customers about what improved, what regressed, and what we do not know yet.
Clad Labs is an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
Apply for this role
* indicates a required field