Staff+ Software Engineer, Capacity Engineering

Anthropic · San Francisco, CA | New York City, NY | Seattle, WA · Software Engineering - Infrastructure · listed July 14, 2026

The shape of it

Seniority
Staff
Where
Not stated
Stated pay
$320,000 – $485,000 USD
Requirements listed
7
Length
1,579 words

In the posting’s own words

As an engineer on Capacity Engineering, you will build the production systems that power this work: data pipelines that ingest and normalize telemetry from heterogeneous cloud environments, observability tooling that gives the org real-time visibility into fleet health, and performance instrumentation that measures how efficiently every major workload uses the hardware it’s running on. You will be expected to write production-quality code every day, operate alongside Kubernetes-native infrastructure at meaningful scale, and directly influence decisions around one of Anthropic’s largest areas of spend.

What it asks for · 7

  • A strong track record building and operating production systems. This is a hands-on engineering role with a devops flavor.
  • Deep experience with at least one major cloud provider (Amazon Web Services, Google Cloud, or Microsoft Azure) and its operations
  • Experience with observability tooling stack, including Prometheus, PromQL, and Grafana, including writing recording rules and building monitoring that engineering teams rely on.
  • Ability to gather your own requirements and work across organizational boundaries in an ambiguous environment with limited direction.
  • Experience with capacity planning, resource management, or cost attribution systems at a hyperscaler or in a large-scale machine learning environment. Time spent in product engineering and developer experience absolutely counts here.
  • Accelerator infrastructure familiarity. GPU metrics (DCGM), TPU utilization, Trainium power and utilization metrics, or experience with machine learning training and inference systems at the hardware level.
  • Experience building internal data products with self-service access, schema contracts, API serving, documentation, and discoverability. Not just pipelines, but thinking about how the data gets consumed.

Also a plus

  • Experience with capacity planning, resource management, or cost attribution systems at a hyperscaler or in a large-scale machine learning environment. Time spent in product engineering and developer experience absolutely counts here.
  • Scheduling and packing efficiency experience, or profiling-driven optimization of large distributed workloads.
  • Multi-cloud data ingestion experience, especially normalizing billing exports, reservation APIs, on-demand capacity reservations, commitments, and vendor telemetry from providers with different billing arrangements.
  • Total cost of ownership and forecasting experience, including decomposing whether infrastructure growth is causal or correlated with business drivers.
  • Accelerator infrastructure familiarity. GPU metrics (DCGM), TPU utilization, Trainium power and utilization metrics, or experience with machine learning training and inference systems at the hardware level.
  • Experience building internal data products with self-service access, schema contracts, API serving, documentation, and discoverability. Not just pipelines, but thinking about how the data gets consumed.
  • Storage efficiency, retention, and lifecycle program experience at exabyte scale.

What the job covers

  • Build the planning and allocation stack — the tools leadership uses to allocate capacity, teams use to plan against their allocations, and the scheduler enforces. Cross-region and cross-provider placement, guardrails, queueing, occupancy KPIs.
  • Drive the efficiency programs: stranding and rightsizing, unused capacity recovery, and job-level utilization across training, inference, and eval. Establish per-config baselines and work with system-owning teams to close the gaps. Utilization improvements are worth enormous sums at our scale.
  • Own attribution and forecasting — reconcile billing across ten-plus providers against telemetry and internal systems, attribute spend to the workloads that generate it, and turn demand signals and research roadmaps into a defensible compute plan and supply pipeline.
  • Build the data platform underneath all of it: pipelines ingesting occupancy, utilization, and cost from a rapidly diversifying fleet into BigQuery, with real ownership of completeness, latency SLOs, and gap detection. Every new provider is a net-new integration.
  • Operate Kubernetes-native systems at scale — collection agents, workload labeling, and the taint/reservation/scheduling behavior that determines what capacity is actually usable.
  • Treat the output as a product, not a pipeline. Gather your own requirements, define schema contracts, and design for consumers ranging from research engineers to a CFO — including on-call and SLOs, because these surfaces are load-bearing for the company.

Tools and skills named

Cloud & infra
  • Kubernetes4×
  • Observability3×
  • AWS
  • Azure
  • GCP
  • Grafana
  • Prometheus
Models & research
  • Inference5×
  • Machine learning2×
  • GPU
Languages
  • Python2×
  • SQL2×
Go to market
  • Forecasting3×
Ways of working
  • On-call
  • Technical writing
Data
  • Data pipelines
Frameworks
  • REST

Words the posting leans on

  • experience10×
  • systems10×
  • data9×
  • engineering9×
  • capacity8×
  • pipeline8×
  • utilization8×
  • infrastructure7×
  • fleet6×
  • workload6×
  • billing5×
  • cloud5×
  • efficiency5×
  • inference5×
  • providers5×
  • research5×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Anthropic

every open role at Anthropic

How this page was made

An automated read of a public job posting, fetched August 25, 2026 and last changed by Anthropic on August 21, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.