Senior Staff Applied AI Engineer - Context Retrieval

Databricks · Mountain View, California; San Francisco, California · Engineering · listed May 7, 2026

The shape of it

Seniority
Staff
Experience asked
10+ years
Where
Hybrid
Stated pay
$228,600 – $342,800 USD
Requirements listed
5
Length
1,375 words

In the posting’s own words

We are hiring a Senior Staff Applied AI Engineer to own context retrieval for Databricks agents across SaaS providers . This is a zero-to-one role with two deeply connected charters:

What it asks for · 5

  • Hands-on experience with modern LLM-era retrieval : RAG architectures, query rewriting, re-ranking with cross-encoders, long-context strategies, and grounding techniques that reduce hallucination.
  • Strong grasp of relevance evaluation : nDCG, MRR, Precision@K, Recall@K; offline/online experimentation; LLM-as-judge frameworks; building human labeling pipelines.
  • Track record of building 0→1 : you've stood up a retrieval system from an empty repo, made the foundational architectural decisions, and grown it into something that customers depend on.
  • Demonstrated ability to operate as a technical leader : setting direction across teams, mentoring senior engineers, and influencing roadmap with research, product, and platform partners.
  • Experience building retrieval over enterprise SaaS sources (permissions, freshness, multi-tenancy, ACL-aware indexing).

Also a plus

  • Experience building retrieval over enterprise SaaS sources (permissions, freshness, multi-tenancy, ACL-aware indexing).
  • Background in agentic systems, tool use, or multi-turn retrieval for LLM agents.
  • Contributions to open-source IR/search projects, or publications at SIGIR, KDD, WWW, EMNLP, or similar venues.
  • Experience training or fine-tuning embedding models, rerankers, or query understanding models.

What the job covers

  • Build the full retrieval stack from scratch. Own the end-to-end system: query understanding, content understanding and indexing, hybrid retrieval, ranking, and evaluation. Make the architectural calls that will define how Databricks agents access context for years to come.
  • Retrieve across heterogeneous data — structured and unstructured. Index and rank across structured assets (tables, columns, SQL queries, dashboards, code, notebooks, jobs) and unstructured content (docs, wikis, tickets, chat, images, video, audio). Each modality has its own signals — design retrieval that exploits them rather than flattens them.
  • Connect to the SaaS surface area customers actually use. Build connectors and retrieval adapters for the systems where enterprise knowledge lives. Treat each retrieval source with its own freshness, permissions, and ranking signals.
  • Optimize for two consumers at once. Retrieval must serve both LLMs (grounded, token-efficient, hallucination-resistant context) and humans (intuitive, explainable discovery). These are different objectives and require different signals — own both.
  • Crack query understanding for agents. Agent queries don't look like web queries. Build query rewriting, decomposition, intent classification, and entity resolution tuned for multi-turn agentic workflows.
  • Crack content understanding at scale. Build the pipelines that extract structure, entities, embeddings, summaries, and metadata from every supported asset type — and keep them fresh as customer data evolves.
  • Build search subagents that reason about retrieval. Design the agentic layer that decides what context is needed , which sources to query , how to decompose and route the search , and — critically — whether the retrieved content is actually sufficient to answer the question . These subagents will plan multi-hop searches, issue follow-up queries when results are weak, ground claims against retrieved evidence, and hand back high-confidence context (or signal failure) to upstream agents. This is where IR meets agentic reasoning.
  • Build the evaluation flywheel for both retrieval and subagents. Stand up offline evals (nDCG, MRR, Recall@K, Precision@K), LLM-as-judge harnesses, human-in-the-loop labeling, and online experimentation. Extend evaluation beyond ranking metrics to measure subagent decision quality — did it ask the right follow-up? , did it correctly recognize when retrieval failed? , did it ground its answer in the right evidence? . Quality you can't measure is quality you can't ship.
  • Set technical direction and grow the team. Set the multi-year roadmap, mentor senior engineers, partner with Research, Product, and Platform leaders, and raise the technical bar across the org.

Tools and skills named

Data
  • Databricks6×
  • Data warehouse2×
  • Experimentation2×
  • Spark2×
  • Elasticsearch
Models & research
  • LLM5×
  • Evaluations
  • Fine-tuning
Go to market
  • SaaS4×
Languages
  • SQL2×
Product & design
  • Roadmap2×
Ways of working
  • Mentorship

Words the posting leans on

  • retrieval26×
  • agents15×
  • build10×
  • systems9×
  • data8×
  • context7×
  • query7×
  • understanding7×
  • agentic6×
  • experience6×
  • building5×
  • customers5×
  • evaluation5×
  • quality5×
  • search5×
  • subagents5×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Databricks

every open role at Databricks

How this page was made

An automated read of a public job posting, fetched August 25, 2026 and last changed by Databricks on August 18, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.