Senior Software Engineer, Chaos Engineering
Datadog · Paris, France · Security · listed September 30, 2026
The shape of it
Seniority
Senior
Where
Hybrid
Requirements listed
6
Length
768 words
In the posting’s own words
At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.
What it asks for · 6
- You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery.
- You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design.
- You have experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important.
- You communicate complex technical decisions clearly through design documents, runbooks, postmortems, and cross-functional technical discussions.
- You are comfortable collaborating across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements.
- Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required.
What the job covers
- Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services.
- Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments.
- Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible.
- Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification.
- Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes.
- Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers.
Tools and skills named
Cloud & infra
- Datadog4×
- Distributed systems2×
- Kubernetes2×
- Observability
Ways of working
- Cross-functional
- Mentorship
- Testing
Frameworks
- gRPC
Product & design
- Design systems
Words the posting leans on
- systems9×
- design7×
- failure7×
- build5×
- engineering5×
- automation4×
- production4×
- reliability4×
- resilience4×
- experience3×
- experiments3×
- failure modes3×
- findings3×
- remediation3×
- services3×
- technical3×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Datadog
- AI Research Engineer - Datadog AI Research (DAIR)Paris, France
- AI Research Scientist - Datadog AI Research (DAIR)Paris, France
- AI Research Scientist - Datadog AI Research (DAIR)New York, New York, USA; Pittsburgh, Pennsylvania, USA
- Applied Science InternParis, France
- Commercial Account ExecutiveSydney, Australia
- Commercial Account ExecutiveDenver, Colorado, USA