Staff Software Engineer - Agent Quality

Databricks · New York City, New York · Engineering - Pipeline · listed May 1, 2026

The shape of it

Seniority
Staff
Experience asked
6+ years
Where
Not stated
Stated pay
$190,000 – $270,000 USD
Requirements listed
6
Length
773 words

In the posting’s own words

As a Staff Software Engineer - Agent Quality, you will be a founding member of a new team focused on evaluating and continuously improving Databricks' AI Agents. You will design and scale the infrastructure, tooling, and developer workflows that let researchers and engineers evaluate agents rigorously — driving a flywheel where evaluation results feed directly back into agent improvement across the full lifecycle, from development and training to production.

What it asks for · 6

  • 6+ years industry experience building software systems
  • Strong Python programming skills, with experience building production or research infrastructure
  • Experience building or operating distributed systems, data pipelines, or large-scale infrastructure with a focus on reliability, correctness, and operational maturity
  • Ability to design pragmatic but rigorous systems that produce trustworthy, reproducible signals for complex applications
  • Comfort working across ambiguous research and product boundaries, and partnering with both researchers and engineers to turn ideas into robust internal platforms
  • A high bar for technical quality, strong ownership, and the ability to influence roadmap and execution across multiple teams

Also a plus

  • Experience with devtools, CI/CD platforms, testing frameworks, observability tooling, or benchmarking infrastructure
  • Familiarity with how LLM or agent quality is measured — whether through evals, experimentation platforms, or production monitoring

What the job covers

  • Stand up the foundational evaluation infrastructure for Genie Agents, enabling rigorous benchmarking, regression detection, and quality measurement across research and product teams.
  • Build the flywheel that connects evaluation results back into agent improvement — closing the loop between production signals, training, and iterative development.
  • Shape the long-term technical direction for agent quality infrastructure, with real influence over how Databricks measures and improves its first-party agents and agent development platform.
  • Help shape the long-term technical direction for agent quality infrastructure as Databricks expands its first-party agents and agent development platform.

Tools and skills named

Data
  • Databricks5×
  • Data pipelines
  • Experimentation
Cloud & infra
  • CI/CD
  • Distributed systems
  • Observability
Models & research
  • Evaluations
  • LLM
Languages
  • Python
Product & design
  • Roadmap
Security & compliance
  • Security
Ways of working
  • Testing

Words the posting leans on

  • agent15×
  • infrastructure7×
  • platform6×
  • agent quality5×
  • development5×
  • data4×
  • engineer4×
  • production4×
  • research4×
  • systems4×
  • experience building3×
  • training3×
  • agent development2×
  • agent improvement2×
  • back2×
  • benchmarking2×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Databricks

every open role at Databricks

How this page was made

An automated read of a public job posting, fetched August 24, 2026 and last changed by Databricks on August 18, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.