Staff Machine Learning Engineer, Public Sector

Scale AI · Denver, CO; Honolulu, HI; Washington, DC · Public Sector Engineering · listed January 30, 2026

The shape of it

Seniority
Staff
Experience asked
8+ years
Where
Not stated
Stated pay
$274,400 – $343,000 USD
Requirements listed
10
Length
1,039 words

In the posting’s own words

The goal of a Staff Machine Learning Engineer at Scale is to lead the design and deployment of agentic AI systems that operate in real-world, mission-critical government environments. On the Public Sector team, you’ll work at the intersection of agentic ML, systems engineering, and applied research, building foundational infrastructure that enables AI systems to reason, plan, and act reliably at national scale.

What it asks for · 10

  • Comfortable with light travel (approximately 10%) for customer interaction and team needs.
  • 8+ years of experience building and deploying applied ML systems in production environments.
  • Deep experience with agentic systems, autonomous workflows, or ML systems that reason and act over multiple steps.
  • Strong background in ML systems engineering, including model serving, pipelines, monitoring, and evaluation.
  • Hands-on experience with retrieval systems, embeddings, or representation learning.
  • Proficiency in Python and modern ML frameworks (ex: PyTorch), with the ability to design systems end to end.
  • Demonstrated ability to operate at Staff-level scope: setting technical direction, owning ambiguous problems, and driving 0→1 initiatives to production.
  • Hands-on experience with geospatial data or GEOINT: reasoning over maps, imagery, or spatial reference systems.
  • Experience building evaluation infrastructure for non-deterministic systems: LLM-as-judge, regression suites for agent behavior, or drift detection in production.
  • A track record of turning a forward-deployed prototype into a supported, documented capability other engineers can deploy without you.

Also a plus

  • Experience deploying ML systems into air-gapped, classified, or otherwise disconnected environments - customer data centers, on-prem infrastructure, or networks with no path to a cloud provider.
  • Prior work with DoD, the intelligence community, or federal mission users - including the judgment to learn a mission well enough to know what "correct" means for the operator using your system.
  • Hands-on experience with geospatial data or GEOINT: reasoning over maps, imagery, or spatial reference systems.
  • Depth in model adaptation - training or fine-tuning embedding models, instruction tuning, LoRA/PEFT, or RLHF.
  • Experience building evaluation infrastructure for non-deterministic systems: LLM-as-judge, regression suites for agent behavior, or drift detection in production.
  • A track record of turning a forward-deployed prototype into a supported, documented capability other engineers can deploy without you.

Tools and skills named

Models & research
  • Machine learning10×
  • Fine-tuning
  • LLM
  • PyTorch
  • Reinforcement learning
Ways of working
  • Testing2×
Languages
  • Python
Product & design
  • Design systems
Security & compliance
  • Security

Words the posting leans on

  • systems19×
  • experience7×
  • production6×
  • agentic systems5×
  • design5×
  • reasoning5×
  • agents4×
  • infrastructure4×
  • model4×
  • data3×
  • engineers3×
  • evaluation3×
  • operate3×
  • public sector3×
  • act2×
  • applied2×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Scale AI

every open role at Scale AI

How this page was made

An automated read of a public job posting, fetched September 10, 2026 and last changed by Scale AI on September 10, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.