Director, Research - AI Evals

Figma · San Francisco, CA • New York, NY • United States · Design · listed July 13, 2026

The shape of it

Seniority
Director
Experience asked
2–10 years
Where
Not stated
Stated pay
$258,000 – $348,000 USD
Requirements listed
6
Length
1,069 words

In the posting’s own words

The Figma Research team is hiring an Director, Research - AI Evals to own how we measure the quality of Figma's AI-powered experiences. As Figma ships more AI capabilities across our products, the question "is this really good?" has never mattered more — and answering it rigorously is what this role exists to do. You'll define what "good" means for our AI features, build the frameworks and quality bars to measure it, and turn that into trusted signal that product teams rely on to decide what to ship.

What it asks for · 6

  • 10+ years of experience in product, research, applied research, or a closely related field, including 2+ years of management experience
  • Direct, hands-on experience owning the evaluation of AI/LLM-powered products
  • Expertise designing and running AI evaluation — human evaluation programs, rubric and benchmark/golden-dataset construction, inter-rater reliability — and sound judgment about when and how to apply automated/model-based approaches (e.g., LLM-as-judge), including their limitations
  • Strength across both qualitative and quantitative methods, comfort with data and metrics, and the ability to reason about model behavior
  • Demonstrated success in identifying the riskiest assumptions behind an ambiguous quality question, prioritizing them, and designing right-sized evaluation to build confidence
  • A proven track record of gaining buy-in from executive and cross-disciplinary stakeholders — transcending methodology to articulate a larger user story and the "so what" to inspire action

Also a plus

  • Experience building or co-building automated evaluation pipelines and regression testing in partnership with engineering, or familiarity with eval tooling (e.g., Braintrust, LangSmith, DeepEval, or equivalents)
  • Experience standing up a new function, practice, or discipline from scratch
  • 2+ years in product design, user-centric product management, data science, product development, and/or front-end engineering
  • A familiarity and depth of experience using Figma's products

What the job covers

  • Own AI evaluation methods and operations for Figma's AI-powered experiences — define quality dimensions, design how we measure them, and turn results into decision-ready signal
  • Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars, combining human evaluation with automated/model-based approaches (e.g., LLM-as-judge) where appropriate
  • Partner with engineering to stand up repeatable, reproducible evaluation pipelines and regression testing, so evaluation is a routine part of how AI features are built and shipped
  • Produce clear readouts and dashboards that let stakeholders confidently make go/no-go and prioritization decisions
  • Socialize a shared definition of quality so evaluation standards are adopted across teams rather than re-invented — and advocate for evaluation as a strategic partner in the product process
  • Manage a small team to execute our AI evals in partnership with contractors, internal staff, and/or LLMs

Tools and skills named

Product & design
  • Figma10×
  • Product management
Models & research
  • LLM5×
  • Evaluations2×
Ways of working
  • Testing2×

Words the posting leans on

  • evaluation14×
  • product13×
  • experience9×
  • quality7×
  • design6×
  • engineering4×
  • research4×
  • build3×
  • evals3×
  • features3×
  • human evaluation3×
  • measure3×
  • ai-powered experiences2×
  • ai/llm-powered products2×
  • and/or2×
  • approaches e.g2×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Figma

every open role at Figma

How this page was made

An automated read of a public job posting, fetched August 25, 2026 and last changed by Figma on July 22, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.