Staff+ Software Engineer, Capacity Engineering
Anthropic · San Francisco, CA | New York City, NY | Seattle, WA · Software Engineering - Infrastructure · listed July 14, 2026
The shape of it
Seniority
Staff
Where
Not stated
Stated pay
$320,000 – $485,000 USD
Requirements listed
7
Length
1,579 words
In the posting’s own words
As an engineer on Capacity Engineering, you will build the production systems that power this work: data pipelines that ingest and normalize telemetry from heterogeneous cloud environments, observability tooling that gives the org real-time visibility into fleet health, and performance instrumentation that measures how efficiently every major workload uses the hardware it’s running on. You will be expected to write production-quality code every day, operate alongside Kubernetes-native infrastructure at meaningful scale, and directly influence decisions around one of Anthropic’s largest areas of spend.
What it asks for · 7
- A strong track record building and operating production systems. This is a hands-on engineering role with a devops flavor.
- Deep experience with at least one major cloud provider (Amazon Web Services, Google Cloud, or Microsoft Azure) and its operations
- Experience with observability tooling stack, including Prometheus, PromQL, and Grafana, including writing recording rules and building monitoring that engineering teams rely on.
- Ability to gather your own requirements and work across organizational boundaries in an ambiguous environment with limited direction.
- Experience with capacity planning, resource management, or cost attribution systems at a hyperscaler or in a large-scale machine learning environment. Time spent in product engineering and developer experience absolutely counts here.
- Accelerator infrastructure familiarity. GPU metrics (DCGM), TPU utilization, Trainium power and utilization metrics, or experience with machine learning training and inference systems at the hardware level.
- Experience building internal data products with self-service access, schema contracts, API serving, documentation, and discoverability. Not just pipelines, but thinking about how the data gets consumed.
Also a plus
- Experience with capacity planning, resource management, or cost attribution systems at a hyperscaler or in a large-scale machine learning environment. Time spent in product engineering and developer experience absolutely counts here.
- Scheduling and packing efficiency experience, or profiling-driven optimization of large distributed workloads.
- Multi-cloud data ingestion experience, especially normalizing billing exports, reservation APIs, on-demand capacity reservations, commitments, and vendor telemetry from providers with different billing arrangements.
- Total cost of ownership and forecasting experience, including decomposing whether infrastructure growth is causal or correlated with business drivers.
- Accelerator infrastructure familiarity. GPU metrics (DCGM), TPU utilization, Trainium power and utilization metrics, or experience with machine learning training and inference systems at the hardware level.
- Experience building internal data products with self-service access, schema contracts, API serving, documentation, and discoverability. Not just pipelines, but thinking about how the data gets consumed.
- Storage efficiency, retention, and lifecycle program experience at exabyte scale.
What the job covers
- Build the planning and allocation stack — the tools leadership uses to allocate capacity, teams use to plan against their allocations, and the scheduler enforces. Cross-region and cross-provider placement, guardrails, queueing, occupancy KPIs.
- Drive the efficiency programs: stranding and rightsizing, unused capacity recovery, and job-level utilization across training, inference, and eval. Establish per-config baselines and work with system-owning teams to close the gaps. Utilization improvements are worth enormous sums at our scale.
- Own attribution and forecasting — reconcile billing across ten-plus providers against telemetry and internal systems, attribute spend to the workloads that generate it, and turn demand signals and research roadmaps into a defensible compute plan and supply pipeline.
- Build the data platform underneath all of it: pipelines ingesting occupancy, utilization, and cost from a rapidly diversifying fleet into BigQuery, with real ownership of completeness, latency SLOs, and gap detection. Every new provider is a net-new integration.
- Operate Kubernetes-native systems at scale — collection agents, workload labeling, and the taint/reservation/scheduling behavior that determines what capacity is actually usable.
- Treat the output as a product, not a pipeline. Gather your own requirements, define schema contracts, and design for consumers ranging from research engineers to a CFO — including on-call and SLOs, because these surfaces are load-bearing for the company.
Tools and skills named
Cloud & infra
- Kubernetes4×
- Observability3×
- AWS
- Azure
- GCP
- Grafana
- Prometheus
Models & research
- Inference5×
- Machine learning2×
- GPU
Languages
- Python2×
- SQL2×
Go to market
- Forecasting3×
Ways of working
- On-call
- Technical writing
Data
- Data pipelines
Frameworks
- REST
Words the posting leans on
- experience10×
- systems10×
- data9×
- engineering9×
- capacity8×
- pipeline8×
- utilization8×
- infrastructure7×
- fleet6×
- workload6×
- billing5×
- cloud5×
- efficiency5×
- inference5×
- providers5×
- research5×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive, Public SectorSydney, Australia
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA