Published Date
3 weeks ago
Work Arrangement
Hybrid • San Francisco, CA
Open Positions
2 openings
Experience Level
Senior
About the opportunity
Advance Constitutional AI, RLHF, and automated prompt red-teaming across Claude model releases.
What you will do
- check_circle Curate high-quality instruction datasets and formulate constitutional alignment critique principles.
- check_circle Train and evaluate Direct Preference Optimization (DPO) and PPO alignment models at scale.
- check_circle Build automated red-teaming agents that probe safety boundaries, prompt injections, and jailbreaks.
- check_circle Collaborate with interpretability researchers to inspect internal neural activations during reasoning.
What we are looking for
- arrow_circle_right Strong track record in deep learning with specific focus on Large Language Model fine-tuning and alignment.
- arrow_circle_right Experience with PyTorch, DeepSpeed, Megatron, DPO, and distributed training clusters.
- arrow_circle_right Passionate dedication to AI safety, truthfulness, and ethical technology development.
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.