Sr. Engineering Manager, AI Runtime
Databricks · Mountain View, California; San Francisco, California · Executive Engineering - Pipeline · listed July 6, 2026
The shape of it
Seniority
Manager
Experience asked
3–8 years
Where
Not stated
Stated pay
$228,600 – $297,120 USD
Requirements listed
9
Length
781 words
In the posting’s own words
As a Senior Engineering Manager, you will lead the team owning both the product experience and the foundational infrastructure of AIR. You'll shape customer-facing capabilities while designing for scalability, extensibility, and performance of GPU training and adjacent areas, collaborating closely across the platform, product, infrastructure, and research organizations.
What it asks for · 9
- 8+ years of software engineering experience, with 3+ years in engineering management.
- Track record building and operating managed GPU training infrastructure at scale (100s/1000s GPUs).
- Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed, Composer, Megatron-LM) and parallelism strategies (FSDP, tensor/pipeline parallelism).
- Experience with training resilience patterns: checkpointing, elastic training, and automated failure recovery for long-running jobs.
- Understanding of GPU performance fundamentals including NCCL, interconnect topologies, and memory optimization.
- Experience building platform products with clear SLAs where you've owned the customer experience, not just the backend.
- Strong cross-functional leadership across platform, product, and research teams, with the ability to lead through ambiguity and deliver complex projects.
- Excellent collaboration and communication skills across engineering, product, and research organizations.
- BS/MS in Computer Science, Electrical Engineering, or related technical field.
What the job covers
- Lead, mentor, and grow a high-performing engineering team responsible for the Custom Training product and its foundational infrastructure, including distributed training orchestration, cluster lifecycle, fault tolerance, and training efficiency.
- Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
- Collaborate closely with product, research, platform, infrastructure teams, and customers to drive end-to-end delivery, from ideation and prioritization to launch and operation.
- Drive architectural decisions and product design for managed GPU training at scale.
- Advocate for customer needs through direct engagement, ensuring engineering decisions translate to clear product impact.
- Build observability and reliability practices for long-running, multi-node training jobs, including checkpoint strategies, failure recovery, and operational runbooks.
- Partner with recruiting to attract, hire, and develop top-tier engineering talent.
Degree language
- BS/MS in Computer Science, Electrical Engineering, or related technical field.
Tools and skills named
Models & research
- GPU6×
- Deep learning
- Fine-tuning
- LLM
- PyTorch
Data
- Databricks2×
Cloud & infra
- Observability
Frameworks
- Node.js
Operations & finance
- Recruiting
Product & design
- Roadmap
Ways of working
- Cross-functional
Words the posting leans on
- training12×
- product11×
- engineering8×
- infrastructure7×
- customer6×
- experience6×
- platform5×
- model4×
- research4×
- air3×
- building3×
- data3×
- deep3×
- gpu training3×
- lead3×
- build2×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.