Sr. Staff Technical Program Manager - Reliability

Databricks · Mountain View, California; San Francisco, California · Engineering - Pipeline · listed March 2, 2026

The shape of it

Seniority
Manager
Experience asked
10+ years
Where
Not stated
Stated pay
$189,800 – $256,160 USD
Requirements listed
9
Length
1,033 words

In the posting’s own words

We are seeking an exceptional Senior Staff Technical Program Manager (TPM) for Reliability to lead the strategy, execution, and continuous improvement of our most critical Reliability initiatives across infrastructure and product engineering teams at Databricks. As Databricks scales to support thousands of customers and the world’s most data-intensive workloads, Reliability is foundational to our mission. In this role, you will lead cross-company programs that significantly enhance the reliability, performance, and operational excellence of our multi-cloud infrastructure.

What it asks for · 9

  • 10+ years of experience managing and delivering large-scale technical programs in cloud infrastructure, distributed systems, SRE, or platform engineering environments.
  • Experience developing infrastructure at two or more hyperscale cloud providers (e.g., AWS, Azure, GCP), with knowledge of cloud primitives, multi-AZ/region architecture, and control plane/data plane patterns.
  • Demonstrated success leading Reliability Programs at scale - including availability, failover, operational excellence, incident reduction, or dependency hardening.
  • Strong understanding of infrastructure, distributed systems, or SRE practices; previous engineering or SRE experience is highly preferred.
  • Experience partnering directly with senior engineering leadership to define strategy and drive large, multi-team initiatives.
  • Ability to translate ambiguous goals into actionable program plans with clear milestones, KPIs, and success metrics.
  • Demonstrated ability to manage complex cross-organizational dependencies, technical risks, and multi-quarter timelines.
  • Experience delivering programs across multiple clouds and/or large-scale cloud-native services.
  • Experience building and scaling engineering processes, operational frameworks, and stakeholder alignment mechanisms.

Also a plus

  • Background in distributed systems engineering, SRE, platform infrastructure, or cloud services.
  • Experience with large-scale compute fleets, container orchestration, autoscaling, or control-plane architecture.
  • Familiarity with reliability methodologies such as SLOs, error budgets, chaos engineering, failure mode analysis, and incident management frameworks.
  • Expertise using Jira or equivalent tools for program tracking and execution.
  • Bachelor’s degree in Computer Science, Engineering, or related technical field; advanced degree preferred.

Tools and skills named

Cloud & infra
  • Site reliability7×
  • Distributed systems5×
  • AWS
  • Azure
  • GCP
Ways of working
  • Cross-functional
  • Jira
  • Technical writing
Data
  • Databricks2×
Security & compliance
  • Risk management
  • Security
Go to market
  • Partnerships
Operations & finance
  • Program management
Product & design
  • Roadmap

Words the posting leans on

  • engineering17×
  • reliability16×
  • programs14×
  • cloud9×
  • experience8×
  • infrastructure8×
  • technical8×
  • operational7×
  • sre7×
  • drive6×
  • distributed systems5×
  • alignment4×
  • define4×
  • execution4×
  • incident4×
  • large-scale4×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Databricks

every open role at Databricks

How this page was made

An automated read of a public job posting, fetched August 25, 2026 and last changed by Databricks on August 18, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.