Senior Product Solutions Architect - Autonomous Incident Response
Datadog · Boston, Massachusetts, USA; Denver, Colorado, USA; New York, New York, USA; San Francisco, California, USA · Product Solutions Architecture · listed August 21, 2026
The shape of it
Seniority
Principal
Experience asked
5+ years
Where
Hybrid
Stated pay
$185,000 – $246,000 USD
Requirements listed
10
Length
1,024 words
In the posting’s own words
The Product Solutions Architecture (PSA) team acts as a technical multiplier across Datadog. PSAs are domain experts who partner with Field teams on complex customer use cases across pre- and post-sales engagements and scale their impact by producing reusable collateral, including reference architectures, technical guides, and enablement assets. By feeding real-world customer insights back to Datadog Product teams, PSAs help influence product roadmaps while accelerating adoption, usage, and long-term customer success.
What it asks for · 10
- You bring a strong engineering foundation, with hands-on experience building and operating incident response programs, on-call rotations, or internal automation/tooling in production environments at a large scale, and are comfortable diving deep into codebases to understand architectural decisions, tradeoffs, and implementation details
- You are familiar with distributed systems and core observability concepts, such as tracing, metrics, and logging, and understand how alerting, event correlation, and incident management fit into a broader operational workflow
- Experience with languages such as Python and/or JavaScript/TypeScript, and familiarity with low-code/no-code application development concepts
- Comfortable operating in rapidly evolving, ambiguous technical domains
- You build deep context across teams and translate it into reusable, scalable solutions
- You take ownership from problem definition through implementation and measurable outcomes
- Highly detail-oriented, particularly when working on architectures and customer-facing technical assets
- You bring strong listening and consultative skills and are experienced supporting customers in high-pressure, complex situations — including live incident scenarios
- Able to sit up to 4 hours, traveling to and from client sites
- Able to travel via auto, train, or air up to 40% of the time
Also a plus
- Successful track record with 5+ years experience working as a Solutions Architect or Sales Engineer supporting incident management, ITSM, on-call/paging, or workflow automation tools (e.g., PagerDuty, Opsgenie, ServiceNow, Jira Service Management, Rundeck)
- Experience using Datadog and/or other observability tools in an SRE or DevOps capacity, particularly owning on-call responsibilities or building internal automation
- Experience building custom applications or internal tools using low-code platforms (e.g., Retool, App Builder) is a plus
What the job covers
- Serve as the in-house subject matter expert for Datadog's Service Management products — App Builder, Incident Management, Event Management, Workflow Automation, and On-Call
- Partner with Field teams to provide hands-on technical and architectural guidance to enterprise customers adopting Service Management, including designing on-call escalation policies, incident response workflows, event correlation pipelines, and custom internal tools with App Builder
- Create high-impact technical collateral, including reference architectures, technical guides, cookbooks, and documentation to enable Field teams and the broader customer community — covering topics like automated incident response, alert noise reduction via Event Management, custom runbook automation with Workflow Automation, and low-code app development with App Builder
- Build proofs of concept and small-scale deployments to validate solutions and reproduce real-world customer environments, such as end-to-end incident lifecycles from alert to resolution
- Act as a trusted advisor to Product Management by delivering actionable feedback informed by real-world field experience, helping shape the roadmap for automation, on-call operations, and incident workflows
Tools and skills named
Cloud & infra
- Datadog8×
- Observability2×
- Distributed systems
- Site reliability
Ways of working
- On-call8×
- Jira
- Technical writing
Go to market
- Solutions architecture3×
- Customer success
Languages
- JavaScript
- Python
- TypeScript
Product & design
- Product management
- Roadmap
Words the posting leans on
- management15×
- incident12×
- customer9×
- on-call7×
- product7×
- technical7×
- app builder6×
- experience6×
- architectures5×
- incident management5×
- service management5×
- solutions5×
- workflow automation5×
- event management4×
- field4×
- incident response4×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Datadog
- AI Research Engineer - Datadog AI Research (DAIR)Paris, France
- AI Research Scientist - Datadog AI Research (DAIR)Paris, France
- AI Research Scientist - Datadog AI Research (DAIR)New York, New York, USA
- Area Vice President, Sales EngineeringBoston, Massachusetts, USA; Denver, Colorado, USA; New York, New York, USA; San Francisco, California, USA
- Associate Field Marketing Manager (US East and Canada)New York, New York, USA
- Associate Growth Marketing ManagerNew York, New York, USA; San Francisco, California, USA