Senior Software Engineer, Chaos Engineering

Datadog · Paris, France · Security · listed September 30, 2026

The shape of it

Seniority
Senior
Where
Hybrid
Requirements listed
6
Length
768 words

In the posting’s own words

At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

What it asks for · 6

  • You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery.
  • You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design.
  • You have experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important.
  • You communicate complex technical decisions clearly through design documents, runbooks, postmortems, and cross-functional technical discussions.
  • You are comfortable collaborating across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements.
  • Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required.

What the job covers

  • Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services.
  • Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments.
  • Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible.
  • Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification.
  • Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes.
  • Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers.

Tools and skills named

Cloud & infra
  • Datadog4×
  • Distributed systems2×
  • Kubernetes2×
  • Observability
Ways of working
  • Cross-functional
  • Mentorship
  • Testing
Frameworks
  • gRPC
Product & design
  • Design systems

Words the posting leans on

  • systems9×
  • design7×
  • failure7×
  • build5×
  • engineering5×
  • automation4×
  • production4×
  • reliability4×
  • resilience4×
  • experience3×
  • experiments3×
  • failure modes3×
  • findings3×
  • remediation3×
  • services3×
  • technical3×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Datadog

every open role at Datadog →

How this page was made

An automated read of a public job posting, fetched September 30, 2026 and last changed by Datadog on September 30, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.