Research Engineer, Model Evaluations
Anthropic · Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY · AI Research & Engineering · listed April 28, 2026
The shape of it
Seniority
Not stated
Where
Hybrid
Stated pay
$500,000 – $850,000 USD
Requirements listed
5
Length
1,195 words
In the posting’s own words
We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do. Your work will turn ambiguous notions of "intelligence" into clear, defensible metrics that researchers, leadership, and the public can rely on.
What it asks for · 5
- Strong Python programming skills, including production or research infrastructure
- Experience building or operating distributed systems, data pipelines, or other infrastructure that needs to be reliable at scale
- Clear written and verbal communication, especially when explaining technical results to non-specialists
- Comfort operating in an on-call or production-support capacity when training runs are live
- Care about the societal impacts of your work and an interest in steering powerful AI to be safe and beneficial
Also a plus
- Hands-on experience using large language models such as Claude, including prompting, sampling, and scaffolding
- Background in data visualization and a track record of building dashboards people actually trust and use
- Experience developing robust evaluation metrics for language models
- Experience with observability, monitoring, or experiment-tracking systems
- Background in statistics and experimental design
- Experience with large-scale dataset sourcing, curation, and processing
- Experience running or supporting ML training infrastructure
- A bias toward picking up slack and operating flexibly across team boundaries
What the job covers
- Design and run new evaluations of Claude's capabilities — reasoning, agentic behavior, knowledge, safety properties — and produce visualizations that make the results legible to researchers and decision-makers
- Build and harden the distributed eval execution platform so hundreds of evals run reliably against checkpoints throughout production RL training runs
- Own the dashboards researchers and leadership use to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions impossible to miss
- Debug anomalous eval results mid-training-run, determine whether the cause is a model change or an infrastructure issue, and communicate the answer clearly under time pressure
- Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations
- Partner with research teams across the full lifecycle of a new capability — from defining what to measure to interpreting results as training progresses
- Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks
- Communicate evaluations and their results to internal stakeholders and, where appropriate, external audiences
Tools and skills named
Models & research
- Evaluations6×
- Prompt engineering2×
- LLM
- Machine learning
Cloud & infra
- Observability2×
- Distributed systems
Data
- Data pipelines
- Statistics
Ways of working
- On-call
- Slack
Languages
- Python
Words the posting leans on
- results8×
- eval7×
- evaluations6×
- experience6×
- infrastructure6×
- researchers6×
- run6×
- training6×
- model5×
- build4×
- capability4×
- claude4×
- research4×
- dashboards3×
- data3×
- design3×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive, Public SectorSydney, Australia
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA