Staff Infrastructure Software Engineer, Enterprise AI
Scale AI · New York, NY; San Francisco, CA · Applications Platform Engineering · listed August 20, 2025
The shape of it
Seniority
Staff
Experience asked
5+ years
Where
Not stated
Stated pay
$216,200 – $270,250 USD
Requirements listed
3
Length
847 words
In the posting’s own words
The ideal candidate thrives in a fast-paced environment, has a passion for both deep technical work and mentoring, and is capable of setting a long-term technical strategy for a critical domain while maintaining a strong, hands-on delivery focus. You will architect and implement solutions across multiple cloud providers (GCP, Azure, AWS) for customers in diverse, highly-regulated industries like healthcare, telecom, finance, and retail.
What it asks for · 3
- Proven experience in a senior role, with 5+ years of full-time software engineering experience.
- Extensive experience with at least one major cloud provider (AWS, Azure, or GCP).
- Proficiency in Python or JavaScript/TypeScript, and SQL.
What the job covers
- Architect multi-cloud systems and abstractions to allow the SGP platform to run on top of existing Cloud providers.
- Use our own data and AI platform to analyze build and test logs and metrics to identify areas for improvement.
- Define the architectural patterns for our multi-cloud infrastructure to support secure, reliable, and scalable Agentic workflows for enterprise customers.
- Enhance engineering and infrastructure efficiency, reliability, accuracy, and response times, including CI/CD processes, test frameworks, data quality assurance, end-to-end reconciliation, and anomaly detection.
- Collaborate with platform and product teams to develop and implement innovative infrastructure that scales to meet evolving needs.
- Design and champion highly scalable, reliable, and low-latency infrastructure and frameworks for building, orchestrating, and evaluating multi-agent systems at enterprise scale.
- Lead the infrastructure roadmap with a strong focus on compliance, privacy, and security standards, including designing change management and data isolation strategies.
- Own the development and maintenance of our best-in-class Agentic observability platform (logging, metrics, tracing, and analytics) to proactively ensure system health and enable rapid incident response.
- Drive developer efficiency by building automated tooling and championing Infrastructure-as-Code (IaC) paradigms throughout the engineering organization to improve workflows and operational efficiency.
Tools and skills named
Cloud & infra
- AWS2×
- Azure2×
- CI/CD2×
- GCP2×
- Observability2×
- Datadog
- Grafana
- Kubernetes
- Prometheus
- Terraform
Languages
- JavaScript
- Python
- SQL
- TypeScript
Security & compliance
- Security2×
Models & research
- LLM
Product & design
- Roadmap
Ways of working
- Mentorship
Words the posting leans on
- infrastructure6×
- platform5×
- data4×
- experience4×
- e.g3×
- efficiency3×
- engineering3×
- enterprise3×
- focus3×
- systems3×
- agentic2×
- architect2×
- aws2×
- azure2×
- building2×
- ci/cd2×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Scale AI
- AI Advisory ConsultantSan Francisco, CA; New York, NY
- AI Advisory PrincipalSan Francisco, CA; New York, NY
- AI Applications Ops Manager, GPSDoha, Qatar
- AI Builder InternSan Francisco, CA
- AI Infrastructure Engineer, Model Serving PlatformSan Francisco, CA; New York, NY
- AI Infrastructure Engineer, Sandbox PlatformLondon, UK