Research Engineer, Pretraining Scaling
Anthropic · San Francisco, CA · AI Research & Engineering · listed September 30, 2025
The shape of it
Seniority
Not stated
Where
Hybrid
Stated pay
$350,000 – $850,000 USD
Requirements listed
8
Length
1,271 words
In the posting’s own words
This role lives at the boundary between research and engineering. You'll work across our entire production training stack: performance optimization, hardware debugging, experimental design, and launch coordination. During launches, the team works in tight lockstep, responding to production issues that can't wait for tomorrow.
What it asks for · 8
- Have hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems
- Genuinely enjoy both research and engineering work—you'd describe your ideal split as roughly 50/50 rather than heavily weighted toward one or the other
- Are excited about being on-call for production systems, working long days during launches, and solving hard problems under pressure
- Thrive when working on whatever is most impactful, even if that changes day-to-day based on what the production model needs
- Excel at debugging complex, ambiguous problems across multiple layers of the stack
- Communicate clearly and collaborate effectively, especially when coordinating across time zones or during high-stress incidents
- Are passionate about the work itself and want to refine your craft as a research engineer
- Care about the societal impacts of AI and responsible scaling
Also a plus
- Previous experience training LLM’s or working extensively with JAX/TPU, PyTorch, or other ML frameworks at scale
- Contributed to open-source LLM frameworks (e.g., open_lm, llm-foundry, mesh-transformer-jax)
- Published research on model training, scaling laws, or ML systems
- Experience with production ML systems, observability tools, or evaluation infrastructure
- Background as a systems engineer, quant, or in other roles requiring both technical depth and operational excellence
What the job covers
- Own critical aspects of our production pretraining pipeline, including model operations, performance optimization, observability, and reliability
- Debug and resolve complex issues across the full stack—from hardware errors and networking to training dynamics and evaluation infrastructure
- Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance
- Respond to on-call incidents during model launches, diagnosing problems quickly and coordinating solutions across teams
- Build and maintain production logging, monitoring dashboards, and evaluation infrastructure
- Add new capabilities to the training codebase, such as long context support or novel architectures
- Collaborate closely with teammates across SF and London, as well as with Tokens, Architectures, and Systems teams
- Contribute to the team's institutional knowledge by documenting systems, debugging approaches, and lessons learned
Tools and skills named
Models & research
- Machine learning5×
- LLM4×
- JAX3×
- PyTorch2×
Cloud & infra
- Observability2×
- Distributed systems
Ways of working
- On-call2×
Operations & finance
- Excel
Words the posting leans on
- model10×
- training10×
- production9×
- systems9×
- research6×
- engineer4×
- experience4×
- launches4×
- performance4×
- build3×
- debugging3×
- engineering3×
- evaluation infrastructure3×
- incidents3×
- issues3×
- operational3×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive, Public SectorSydney, Australia
- Account Executive - Public Sector (ASEAN)Singapore
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA
- AI Fluency Education LeadSan Francisco, CA | New York City, NY