Manager, Software Engineering - Observability

Figma · San Francisco, CA • New York, NY • United States · Engineering · listed February 27, 2026

The shape of it

Seniority
Manager
Experience asked
4+ years
Where
Not stated
Stated pay
$258,000 – $376,000 USD
Requirements listed
5
Length
1,196 words

In the posting’s own words

This team ensures that engineers across Figma can detect issues quickly, understand system behavior at scale, and make informed decisions about reliability. The team owns and evolves our core observability stack—including platforms like Datadog, shared instrumentation libraries, the agents and operators that power telemetry collection, and a host of internally developed components to power AI Trace Observability —while continuously raising the bar on signal quality and operational clarity.

What it asks for · 5

  • 4+ years of experience leading infrastructure, observability, or platform engineering teams, with a track record of delivering highly reliable production systems
  • Deep hands-on experience with modern observability platforms (e.g., Datadog, OpenTelemetry) across metrics, logs, and distributed tracing
  • Strong understanding of distributed systems, instrumentation best practices, SLO design, and incident response workflows
  • Experience driving cost transparency and accountability initiatives, including cost attribution, budgeting, forecasting, and alerting in cloud environments
  • Demonstrated ability to set technical direction, drive cross-functional alignment (Engineering, Finance, Security), and make sound architectural decisions in complex environments

Also a plus

  • Experience building observability, telemetry, data infrastructure, or evaluation systems for AI or machine learning products, including familiarity with LLM traces, prompt and model versioning, classifiers, and offline or online eval workflows.
  • Experience designing or evolving company-wide observability standards, shared libraries, and agent/operator-based integrations.
  • Background in cost optimization for infrastructure or observability tooling, including vendor negotiations and usage modeling.
  • Experience applying AI or machine learning techniques to anomaly detection, root cause analysis, or operational automation.
  • Familiarity with OpenTelemetry and modern instrumentation frameworks across multiple programming languages.
  • Experience scaling and mentoring high-performing engineering teams through platform expansion or significant architectural change.

What the job covers

  • Lead and grow a team of 5 engineers responsible for the reliability, scalability, and evolution of Figma’s Observability and AI Trace Observability platforms.
  • Own and operate the AI Trace Observability ecosystem, including a safe telemetry pipeline that aggregates all relevant data, runs it through classification and then gives engineers the ability to run evals against this data.
  • Own and operate Figma’s core observability stack, including vendor platforms such as Datadog, ensuring high availability, strong data quality, and effective signal-to-noise across metrics, logs, and traces.
  • Define and drive the technical strategy for instrumentation standards, observability libraries, agents, and operators used to monitor internal and external facing services.
  • Explore and implement innovative, AI-driven approaches to anomaly detection, root cause analysis, signal correlation, and operational automation.
  • Partner with infrastructure, product engineering, finance, and security teams to improve visibility into system health and cost efficiency at scale.
  • Coach and mentor engineers through career development, performance feedback, and technical leadership, fostering a culture of ownership, collaboration, and high quality execution.

Tools and skills named

Cloud & infra
  • Observability17×
  • Datadog4×
  • Distributed systems2×
Product & design
  • Figma10×
Models & research
  • Machine learning2×
  • Evaluations
  • LLM
Ways of working
  • Cross-functional2×
  • Mentorship
Operations & finance
  • Budgeting2×
Security & compliance
  • Security2×
Go to market
  • Forecasting

Words the posting leans on

  • observability17×
  • platform10×
  • systems9×
  • experience8×
  • engineering7×
  • cost6×
  • engineers6×
  • instrumentation6×
  • trace6×
  • data4×
  • datadog4×
  • design4×
  • infrastructure4×
  • libraries4×
  • operate4×
  • quality4×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Figma

every open role at Figma

How this page was made

An automated read of a public job posting, fetched August 25, 2026 and last changed by Figma on July 23, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.