Senior Applied Scientist - AI Platform
Datadog · Paris, France · Dev Eng · listed September 2, 2026
The shape of it
Seniority
Senior
Experience asked
6+ years
Where
Hybrid
Requirements listed
8
Length
1,117 words
In the posting’s own words
AI Platform builds the foundations of Datadog's AI efforts. The org is 70+ people organised in three pillars: training and serving (GPU clusters, distributed training, low-level infrastructure), agents (agent harnesses, memory systems, the internal AI gateway that routes every LLM request at Datadog), and evaluation and experimentation. This role sits in the evaluation and experimentation pillar, which owns Datadog's shared annotation and evaluation infrastructure — including the evaluation scenario store and the telemetry archival systems used across the Bits org. Together they let an agent travel back in time and query what Datadog looked like at the exact moment an incident happened, so scenarios can be replayed and agent performance tracked over time.
What it asks for · 8
- You have a PhD, MS or equivalent research experience in a scientific field, with strong applied mathematics grounding.
- 6+ years of relevant applied science or ML engineering experience, including setting technical direction for others.
- You have hands-on experience with LLM and agent post-training data: how it is created, managed, and how training-data quality is controlled. This is the requirement that matters most.
- You have real domain expertise in LLMs and agentic applications — not classical ML fine-tuning. Fine-tuning classifiers or traditional models is a different problem from the one this team is solving.
- You have evaluated agents or LLM applications, and can define what 'good' means before you measure it.
- You are a strong programmer and production software engineer. Python at minimum, plus the ability to ship scalable production systems and work with distributed systems.
- You collaborate well across engineering and science teams, and you're comfortable being the domain expert who decides what comes next.
- You thrive in ambiguity and can make sound technical calls when the path isn't yet defined.
Also a plus
- Hands-on LLM fine-tuning, post-training or model training experience.
- Background in statistics, experiment design and data analysis.
- Experience deploying production-level ML infrastructure.
- Observability or monitoring systems background.
- Architecture-level understanding of LLMs.
What the job covers
- Own the applied science direction for GenSim: set the methodology and the forward-looking technical calls on how simulated environments and post-training data should be built, on a team where that decision-making does not exist yet.
- Define, measure and raise the quality of post-training data — basic correctness, representativeness against the real distribution of customer systems and production telemetry, and difficulty — and make those measures something the team can act on release over release.
- Close the realism gap. Simulated environments today are too clean and the injected problems are not yet hard enough; you'll drive the research and the engineering that make them look like real, imperfect production systems.
- Build scalable, production-grade systems rather than research scripts. The output is not just a dataset — it is a system of synthetic environments that must be reliable and invokable inside a training loop.
- Determine how this data is best applied, in LLM post-training and in evaluation, and own the agent and LLM application evaluation approaches for these environments.
- Work cross-functionally with the engineers and applied scientists on adjacent teams — Bits AI SRE, the model training effort, and the wider evaluation and experimentation pillar — so that what you learn moves freely in both directions.
Degree language
- You have a PhD, MS or equivalent research experience in a scientific field, with strong applied mathematics grounding.
Tools and skills named
Models & research
- LLM8×
- Fine-tuning3×
- Machine learning3×
- GPU
Cloud & infra
- Datadog8×
- Site reliability4×
- Distributed systems
- Observability
Data
- Experimentation3×
- Statistics
Languages
- Python
Words the posting leans on
- agent11×
- systems9×
- llm8×
- applied7×
- evaluation7×
- post-training data6×
- experience5×
- real5×
- training5×
- applications4×
- build4×
- direction4×
- gensim4×
- measure4×
- model4×
- problems4×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Datadog
- AI Research Engineer - Datadog AI Research (DAIR)Paris, France
- AI Research Scientist - Datadog AI Research (DAIR)Paris, France
- AI Research Scientist - Datadog AI Research (DAIR)New York, New York, USA; Pittsburgh, Pennsylvania, USA
- Area Vice President, Sales EngineeringBoston, Massachusetts, USA; Denver, Colorado, USA; New York, New York, USA; San Francisco, California, USA
- Commercial Account ExecutiveTokyo, Japan
- Commercial Account ExecutiveSydney, Australia