Senior Software Engineer - Incident Insights & Readiness
Datadog · Boston, Massachusetts, USA; New York, New York, USA · Security · listed October 1, 2026
The shape of it
Seniority
Senior
Experience asked
5+ years
Where
Hybrid
Stated pay
$192,000 – $240,000 USD
Requirements listed
8
Length
1,451 words
In the posting’s own words
We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way
What it asks for · 8
- At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of TypeScript.
- Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes.
- Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution.
- Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings.
- Experience participating in on-call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus.
- Empathy, collaboration, and communication skills in English to cultivate strong relationships across various teams in the organization
- Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority.
- We welcome candidates from a variety of backgrounds, including software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response.
What the job covers
- Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation.
- Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across Datadog. Our aim is to fully support our incident responders in dealing with complexity.
- Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group.
- Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people.
- Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices.
- Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers.
- Lead cross-functional initiatives in engineering organizations across Datadog, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence.
Tools and skills named
Cloud & infra
- Datadog10×
- Site reliability3×
- Distributed systems2×
- Kubernetes2×
Ways of working
- On-call8×
- Cross-functional4×
- Mentorship4×
Languages
- Python2×
- TypeScript2×
Words the posting leans on
- incident27×
- experience16×
- engineering14×
- learning11×
- building8×
- design8×
- on-call8×
- organization8×
- technical8×
- incident response7×
- operational7×
- software7×
- engineers6×
- improve6×
- practices6×
- support6×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Datadog
- AI Research Engineer - Datadog AI Research (DAIR)Paris, France
- AI Research Scientist - Datadog AI Research (DAIR)New York, New York, USA; Pittsburgh, Pennsylvania, USA
- AI Research Scientist - Datadog AI Research (DAIR)Paris, France
- Applied Science InternParis, France
- Commercial Account ExecutiveSeoul, South Korea
- Commercial Account ExecutiveTokyo, Japan