Staff + Senior Software Engineer, Inference

Anthropic · Ontario, CAN · Software Engineering - Infrastructure · listed August 17, 2026

The shape of it

Seniority
Staff
Where
Hybrid
Requirements listed
6
Length
1,303 words

In the posting’s own words

Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry’s largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.

What it asks for · 6

  • Significant software engineering experience, particularly with distributed systems
  • Results-oriented, with a bias towards flexibility and impact
  • Willingness to pick up slack, even if it goes outside your job description
  • Desire to learn more about machine learning systems and infrastructure
  • Thrive in environments where technical excellence directly drives both business results and research breakthroughs
  • Care about the societal impacts of your work

Also a plus

  • Experience with high-performance, large-scale distributed systems
  • Experience implementing and deploying machine learning systems at scale
  • Experience with load balancing, request routing, or traffic management systems
  • Familiarity with LLM inference optimization, batching, and caching strategies
  • Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure)
  • Proficiency in Python or Rust

What the job covers

  • Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide
  • Develop resilient, flexible systems that adapt in real time to real world events
  • Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators
  • Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads
  • Build and operate production-grade deployment pipelines for releasing new models to users
  • Provide high-performance inference infrastructure that enables researchers to develop next-generation models
  • Integrate new AI accelerator platforms and support inference for new model architectures

Tools and skills named

Models & research
  • Inference16×
  • Machine learning4×
  • LLM2×
Cloud & infra
  • Distributed systems8×
  • AWS2×
  • Azure2×
  • GCP2×
  • Kubernetes2×
  • Observability2×
Languages
  • Python2×
  • Rust2×
Ways of working
  • Slack2×

Words the posting leans on

  • systems21×
  • inference16×
  • models12×
  • experience10×
  • routing10×
  • accelerators8×
  • distributed systems8×
  • deployment7×
  • develop7×
  • infrastructure7×
  • research7×
  • fleet5×
  • millions users5×
  • request routing5×
  • autoscaling4×
  • build4×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Anthropic

every open role at Anthropic

How this page was made

An automated read of a public job posting, fetched August 25, 2026 and last changed by Anthropic on August 21, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.