Research Engineer, Code RL (Reinforcement Learning)
Anthropic · San Francisco, CA | New York City, NY · AI Research & Engineering · listed June 11, 2026
The shape of it
Seniority
Not stated
Where
Hybrid
Stated pay
$500,000 – $850,000 USD
Requirements listed
5
Length
1,328 words
In the posting’s own words
We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to write, edit, test, debug, and ship real software — end to end, on real codebases, with real tools — and to do it correctly, fast, and safely.
What it asks for · 5
- Have strong software-engineering skills and deep Python expertise, including async/concurrent programming
- Are comfortable owning systems end to end and debugging across the stack
- Can balance research exploration with engineering implementation, and engage rigorously in shaping experimental design and interpreting results
- Care about code quality, testing, and performance
- Are passionate about the potential impact of AI and are committed to developing safe and beneficial systems
Also a plus
- Experience with reinforcement learning, RLHF, post-training, or LLM finetuning
- Built coding agents, code-execution sandboxes, eval harnesses, verifiers, or developer tooling
- Background in program analysis, testing, verification, compilers, or formal methods
- Experience with PyTorch and large-scale distributed training; performance profiling and optimization of ML systems
- CUDA / GPU or TPU kernel experience and accelerator-performance intuition
- Experience with virtualization and sandboxed code execution environments
Tools and skills named
Models & research
- Reinforcement learning8×
- CUDA2×
- Fine-tuning2×
- GPU2×
- LLM2×
- Machine learning2×
- PyTorch2×
Ways of working
- Testing4×
Languages
- Python2×
Security & compliance
- Security2×
Words the posting leans on
- code11×
- experience8×
- research engineer7×
- coding6×
- end6×
- performance6×
- reinforcement learning6×
- systems6×
- training5×
- engineering4×
- testing4×
- areas3×
- fast code3×
- models3×
- real3×
- accelerator-performance intuition2×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive, Public SectorSydney, Australia
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA