Published Date
1 week ago
Work Arrangement
Remote • Remote — USA
Open Positions
2 openings
Experience Level
Mid-level
About the opportunity
Curate datasets, design evaluation suites, and evaluate model alignment and behavioral benchmarks for Claude.
Key Focus Areas:
• Constitutional AI benchmark verification and red-teaming datasets.
• Quantitative evaluation pipelines measuring reasoning accuracy and refusal rates.
• Tool-use and multi-step agent trajectory evaluations.
• Statistical analysis of model drift and calibration across iterations.
What you will do
- check_circle Create and curate high-quality benchmark datasets for reasoning and alignment evaluations.
- check_circle Write Python scripts to automate prompt batching, model scoring, and statistical drift detection.
- check_circle Analyze qualitative and quantitative failure modes in frontier model generations.
- check_circle Collaborate with alignment researchers to iterate on constitutional AI guidelines.
What we are looking for
- arrow_circle_right Strong Python scripting skills and fluency with pandas, NumPy, and JSON data transformations.
- arrow_circle_right Experience designing prompt templates, automated eval scripts, and model red-teaming tasks.
- arrow_circle_right Analytical mindset with an ability to detect subtle logical or factual inaccuracies in model outputs.
- arrow_circle_right Clear documentation and technical writing abilities.
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.