Staff Research Engineer, Multi-Agent Scaling
Anthropic · San Francisco, CA | New York City, NY | Seattle, WA · AI Research & Engineering · listed October 2, 2026
The shape of it
Seniority
Staff
Where
Hybrid
Stated pay
$500,000 – $850,000 USD
Requirements listed
8
Length
1,164 words
In the posting’s own words
This role lives at the boundary between research and engineering. It is a generalist role on a small team: you'll design and run large experiments, build the systems they run on, and get to the bottom of surprising results. We often need to go from a vague question to a running experiment quickly.
What it asks for · 8
- Have significant software engineering, ML or research engineering experience
- Have owned something substantial end to end, such as a large system, an evaluation or benchmark, an agent product, or a research project
- Genuinely enjoy both research and engineering work
- Think quantitatively about complex systems, and think twice before trusting a number
- Can work from a vague question rather than a spec
- Are results-oriented, with a bias towards flexibility and impact
- Have clear written and verbal communication
- Care about the societal impacts of your work
Also a plus
- Experience building or operating large-scale distributed systems, such as schedulers, sandboxed code execution, or inference and RL infrastructure
- Built evaluations, benchmarks or harnesses for LLMs or agents
- Experience building complex agentic systems that use LLMs
- Experience with scaling laws or other large-scale empirical research
- A background in operations research, statistics, economics, physics, quantitative finance, or another field that models and optimizes complex systems
- Formal certifications or education credentials
- Academic research experience or publication history
- Prior experience with multi-agent systems or reinforcement learning
What the job covers
- Design, run and interpret large-scale experiments on agent teams, reasoning rigorously about what the data does and doesn't show
- Investigate how performance and efficiency change as team size, compute and task horizon grow, and find the bottlenecks that limit them
- Build and scale the systems that run very large agent teams reliably, and debug the failures that only appear at scale
- Design evaluations for long-horizon problems, and keep their results trustworthy
- Build the tooling and metrics that let researchers see what a large agent team is doing and why
- Partner with research teams across Anthropic so they can run their own experiments on the platform, and communicate findings clearly
Tools and skills named
Models & research
- Evaluations3×
- LLM2×
- Inference
- Machine learning
- Reinforcement learning
Cloud & infra
- Distributed systems
Data
- Statistics
Words the posting leans on
- agent13×
- research9×
- systems8×
- large7×
- run7×
- build6×
- experience6×
- evaluations5×
- scale5×
- design4×
- engineering4×
- experiments4×
- large agent4×
- problem4×
- grow3×
- large-scale3×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsDublin, IE
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- AI Deployment Specialist, Beneficial DeploymentsSan Francisco, CA | New York City, NY