Skip to main content
Applied AI & Data Specialist - Model Alignment & Evaluation — Anthropic
Anthropic logo
Generative & Agentic AI · Verified Opening #52

Applied AI & Data Specialist - Model Alignment & Evaluation

business Anthropic location_on Remote — USA (Remote) home_work Remote Full Time

Compensation

$95K – 122K/yr

bolt Apply for this role
calendar_month

Published Date

1 week ago

home_work

Work Arrangement

Remote • Remote — USA

group_add

Open Positions

2 openings

workspace_premium

Experience Level

Mid-level

overview Role Overview

About the opportunity

Curate datasets, design evaluation suites, and evaluate model alignment and behavioral benchmarks for Claude.

Anthropic is looking for an Applied AI & Data Specialist to support alignment, safety, and reasoning evaluation. You will develop synthetic evaluation datasets, measure model responses across diverse safety and domain tasks, and assist in fine-tuning runs. This is a high-impact role helping make frontier models helpful, honest, and harmless.

Key Focus Areas:
• Constitutional AI benchmark verification and red-teaming datasets.
• Quantitative evaluation pipelines measuring reasoning accuracy and refusal rates.
• Tool-use and multi-step agent trajectory evaluations.
• Statistical analysis of model drift and calibration across iterations.
task Core Responsibilities

What you will do

  • check_circle Create and curate high-quality benchmark datasets for reasoning and alignment evaluations.
  • check_circle Write Python scripts to automate prompt batching, model scoring, and statistical drift detection.
  • check_circle Analyze qualitative and quantitative failure modes in frontier model generations.
  • check_circle Collaborate with alignment researchers to iterate on constitutional AI guidelines.
verified_user Candidate Profile

What we are looking for

  • arrow_circle_right Strong Python scripting skills and fluency with pandas, NumPy, and JSON data transformations.
  • arrow_circle_right Experience designing prompt templates, automated eval scripts, and model red-teaming tasks.
  • arrow_circle_right Analytical mindset with an ability to detect subtle logical or factual inaccuracies in model outputs.
  • arrow_circle_right Clear documentation and technical writing abilities.
code_blocks Technologies & Competencies

Skills & Tech Stack

Python Model Evaluation LLM Alignment Prompt Engineering Synthetic Data Statistical Analysis
handshake Direct Referral Network

Why candidate applications stand out

verified

Verified Technical Credentials

Applications include direct proof-of-work repositories and instructor verification endorsements.

person_check

Fast-Track Hiring Visibility

Direct internal referral channels through enterprise partners bypass automated resume discard filters.

work Planning your next career move? Get professional course guidance from TutorAI.

cookie

We value your privacy

We use cookies to run the site and, with your consent, to understand how it's used so we can improve it. See our Cookie Policy and Privacy Policy.