Research Engineer, RL Scaling Science
Anthropic · London, UK · AI Research & Engineering · listed June 22, 2026
The shape of it
Seniority
Not stated
Where
Hybrid
Requirements listed
5
Length
988 words
In the posting’s own words
This role lives at the boundary between research and engineering. The problems are open, the experiments run at frontier scale, and the path from a robust result to production is short.
What it asks for · 5
- Strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area
- Demonstrated ability to own large experiments end-to-end, from design through interpretation
- Proficiency in Python and experience working with large-scale or distributed ML systems
- Comfort operating at the research/systems boundary, including debugging where the two meet
- Care about the societal impacts of AI and responsible scaling
Also a plus
- Published or shipped work in long-horizon RL or RL fundamentals
- Experience translating research findings into production training recipes
- Demonstrated large scale industry impact via RL interventions
- Experience working on frontier-scale training runs with long trajectories
What the job covers
- Design, run, and interpret large-scale RL experiments, reasoning rigorously about what the data does and doesn't show
- Investigate how RL improves as horizon, compute, and model size grow
- Build and maintain benchmarks for long-horizon RL so progress is measurable and reproducible
- Translate validated findings into production training recipes, exercising judgment about when a result is robust enough to ship
- Debug complex issues at the seam where research meets infrastructure - failures that only appear at scale
- Partner closely with adjacent RL teams across research and engineering and advance our overall RL stack
Tools and skills named
Models & research
- Machine learning2×
- Reinforcement learning2×
Languages
- Python
Words the posting leans on
- training7×
- research6×
- run5×
- scale5×
- design4×
- experiments4×
- findings4×
- large-scale4×
- long-horizon4×
- model4×
- benchmarks3×
- experience3×
- production training3×
- scaling3×
- training recipes3×
- adjacent2×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive, Public SectorSydney, Australia
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA