AI Research Engineer - Datadog AI Research (DAIR)

Datadog · Paris, France · Dev Eng · listed August 27, 2025

The shape of it

Seniority
Not stated
Where
Not stated
Requirements listed
6
Length
1,103 words

In the posting’s own words

As a Research Engineer on our team, you will partner with Research Scientists to turn research ideas into working systems, building the data, tooling, and infrastructure that enable rapid iteration, trustworthy evaluation, and a smooth path from prototype to production.

What it asks for · 6

  • You have depth in distributed computing, RL Infra, and ML systems for training and inference at scale; experience with Ray, Slurm, or similar frameworks is a plus
  • You are proficient in Python, familiar with a systems language (e.g., Rust, C++, or Go), and comfortable with modern cloud and data infrastructure
  • You have practical experience implementing and operating ML training and inference systems (e.g., PyTorch or JAX), including containerization, orchestration, and GPU acceleration
  • You have practical experience with large-scale model training and fine-tuning, including frameworks like Megatron-LM, DeepSpeed, SkyRL, VeRL, or TorchTitan, and techniques such as SFT, RLVR, RLHF, and efficient inference (quantization, speculative decoding)
  • You can explain design and performance trade-offs clearly to both technical and non-technical audiences
  • You have experience supporting or contributing to research publications

Also a plus

  • You have strong software engineering skills with experience in domains such as observability, SRE, or security
  • You have experience bridging research prototypes and real-world product applications, especially with large foundation models, world models, or RL-trained agents
  • You have a passion for pushing the boundaries of AI with a focus on customer impact and scalable deployment
  • You have hands-on experience with GPU programming and optimization, including CUDA
  • You have experience writing production data pipelines and applications
  • You have experience building simulation or sandbox environments for agent training

What the job covers

  • Build and operate multimodal data pipelines, training and evaluation infrastructure, benchmarks, and internal tooling
  • Implement models, run experiments at scale, and profile for reliability, performance, and cost
  • Build simulation environments and replay infrastructure for agent training and evaluation
  • Orchestrate distributed training and distributed RL with Ray, including scheduling, scaling, and failure recovery
  • Establish rigorous automated benchmarks and regression tests for world model predictions, agent performance, and simulation fidelity
  • Collaborate with Research Scientists, Product, and Engineering to integrate capabilities into Datadog's products and to harden prototypes into reliable services
  • Contribute to research publications at top-tier conferences (e.g., NeurIPS, ICLR, ICML), and produce high-quality code, documentation, and open-source artifacts

Tools and skills named

Models & research
  • Inference3×
  • GPU2×
  • Machine learning2×
  • CUDA
  • Fine-tuning
  • JAX
  • PyTorch
  • Reinforcement learning
Cloud & infra
  • Datadog4×
  • Observability4×
  • Site reliability2×
  • Distributed systems
Languages
  • C++
  • Python
  • Rust
Security & compliance
  • Security3×
Data
  • Data pipelines2×
Go to market
  • Forecasting
Ways of working
  • Technical writing

Words the posting leans on

  • models11×
  • experience9×
  • research9×
  • training9×
  • agents7×
  • infrastructure6×
  • simulation5×
  • systems5×
  • data4×
  • distributed4×
  • e.g4×
  • evaluation4×
  • observability4×
  • build3×
  • building3×
  • inference3×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Datadog

every open role at Datadog

How this page was made

An automated read of a public job posting, fetched August 25, 2026 and last changed by Datadog on August 24, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.