Staff + Senior Software Engineer, Inference Deployment
Anthropic · San Francisco, CA | New York City, NY | Seattle, WA · Software Engineering - Infrastructure · listed June 29, 2026
The shape of it
Seniority
Staff
Where
Hybrid
Stated pay
$320,000 – $485,000 USD
Requirements listed
6
Length
1,070 words
In the posting’s own words
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry’s largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.
What it asks for · 6
- Significant software engineering experience, particularly with distributed systems
- Results-oriented, with a bias towards flexibility and impact
- Willingness to pick up slack, even if it goes outside your job description
- Desire to learn more about machine learning systems and infrastructure
- Thrive in environments where technical excellence directly drives both business results and research breakthroughs
- Care about the societal impacts of your work
Also a plus
- Experience with high-performance, large-scale distributed systems
- Experience implementing and deploying machine learning systems at scale
- Experience with load balancing, request routing, or traffic management systems
- Familiarity with LLM inference optimization, batching, and caching strategies
- Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure)
- Proficiency in Python or Rust
What the job covers
- Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide
- Develop resilient, flexible systems that adapt in real time to real world events
- Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators
- Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads
- Build and operate production-grade deployment pipelines for releasing new models to users
- Provide high-performance inference infrastructure that enables researchers to develop next-generation models
- Integrate new AI accelerator platforms and support inference for new model architectures
Tools and skills named
Models & research
- Inference11×
- Machine learning2×
- LLM
Cloud & infra
- Distributed systems5×
- AWS
- Azure
- GCP
- Kubernetes
- Observability
Languages
- Python
- Rust
Ways of working
- Slack
Words the posting leans on
- systems13×
- inference11×
- models7×
- routing6×
- accelerators5×
- distributed systems5×
- experience5×
- deployment4×
- develop4×
- infrastructure4×
- research4×
- serve4×
- customers3×
- fleet3×
- millions users3×
- request routing3×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive, Public SectorSydney, Australia
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA