Sr Platform Monitoring Engineer
Databricks · United States · Support · listed August 26, 2026
The shape of it
Seniority
Senior
Experience asked
6+ years
Where
Not stated
Stated pay
$144,600 – $198,900 USD
Requirements listed
6
Length
1,001 words
In the posting’s own words
We are seeking an experienced Tech Lead to shape the future of platform observability and proactive monitoring. This is a high-impact technical role for someone who thrives at the intersection of platform reliability, incident response, and customer obsession.You will serve as a critical first responder for the Databricks Platform, leading complex investigations, designing observability solutions, and driving systemic improvements that enhance customer experience and platform stability.
What it asks for · 6
- Minimum of 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or similar role.
- Production-level experience with at least one major cloud provider (AWS, Azure, GCP) and proficiency in container and orchestration technologies (Docker, Kubernetes).
- Hands-on experience with monitoring, logging, and alerting tools such as ELK, Prometheus, Grafana, PagerDuty, etc. Ability to architect monitoring solutions that correlate metrics, logs, and traces.
- Strong proficiency in Python or similar languages with the ability to build production-quality automation tools.
- Experience owning critical phases of the incident lifecycle from detection through resolution and post-mortem analysis in demanding production environments.
- BS or Master's, or PhD in Computer Science or Computer Engineering, or related Engineering field.
What the job covers
- Lead platform incident investigation, coordinating cross-functional teams through rapid detection, mitigation, and resolution to minimize customer impact.
- Conduct thorough post-incident root cause analysis across infrastructure, services, and cloud providers to identify systemic patterns and prevent future occurrences.
- Design and implement customer-focused alerting pipelines and end-to-end observability workflows to enhance detection coverage and reduce mean time to detection.
- Build automation tools, establish reusable monitoring patterns, and resolve reliability gaps that directly impact customer experience.
- Provide mentorship to junior engineers on observability patterns, alert design, and service health metrics.
- Participate in on-call rotation
Degree language
- BS or Master's, or PhD in Computer Science or Computer Engineering, or related Engineering field.
Tools and skills named
Cloud & infra
- Observability6×
- AWS2×
- Azure2×
- Docker2×
- GCP2×
- Grafana2×
- Kubernetes2×
- Prometheus2×
- Site reliability2×
Ways of working
- Cross-functional3×
- Mentorship2×
- On-call2×
Data
- Databricks3×
Languages
- Python2×
Product & design
- User experience
Words the posting leans on
- experience11×
- detection9×
- monitoring8×
- platform8×
- customer7×
- engineer7×
- incident7×
- tools7×
- observability6×
- patterns6×
- alerting5×
- engineering5×
- impact5×
- services5×
- analysis4×
- automation tools4×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.