Software Engineer, Data Infrastructure

Scale AI · New York, NY; Washington, DC · Public Sector Engineering · listed June 22, 2026

The shape of it

Seniority
Senior
Experience asked
5+ years
Where
On site
Stated pay
$186,400 – $233,000 USD
Requirements listed
4
Length
1,098 words

In the posting’s own words

PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants.

What it asks for · 4

  • Engineering Excellence: Deep, expert-level proficiency in systems languages (e.g., Rust, Go, C++, or highly optimized Python/Java, Spark, PySpark) and a fundamental understanding of memory management, compute limits, and distributed systems architecture.
  • High-Throughput / Low-Latency Data: Proven track record of processing massive datasets. You understand how to optimize massive batch jobs and parallel processing across distributed simulation nodes without sacrificing speed.
  • Experience with LLM context optimization, vector embeddings, or agentic AI frameworks (e.g., advanced RAG architectures).
  • Previous experience in a high-growth, 0-to-1 startup environment.

Also a plus

  • Security Clearance: An active Secret or TS/SCI clearance is a nice to have for this role. If you do not have an active clearance, you must be eligible and willing to obtain one.
  • Experience with LLM context optimization, vector embeddings, or agentic AI frameworks (e.g., advanced RAG architectures).
  • Deep domain experience working with wargaming data, complex systems modeling, or distributed simulation protocols.
  • Previous experience in a high-growth, 0-to-1 startup environment.

What the job covers

  • Architect the Data Ensemble: Design and implement the architecture to ensemble various sources of injected context (deeply structural simulation data, historical game states, and dynamic user inputs) into a unified, highly queryable format optimized for LLM consumption.
  • Massive Batch Infrastructure: Build highly scalable, resilient data architectures from scratch. You will optimize for moving, transforming, and processing massive quantities of simulation output data via enormous batch jobs, maintaining the minimal latency required for rapid wargame iterations.
  • Complex Data Modeling: Design sophisticated, highly relational data models that accurately represent massive, state-based simulation environments, making them easily interpretable by machine learning models.
  • First-Principles Problem Solving: Navigate highly ambiguous product requirements to design custom, ground-up systems where existing open-source or enterprise tools simply cannot handle the structural complexity or scale.
  • Technical Leadership: Set the technical standard for the data infrastructure team, driving rigorous code quality, system performance, and architectural clarity.

Tools and skills named

Languages
  • C++
  • Go
  • Java
  • Python
  • Rust
Data
  • Spark2×
  • Data modeling
Models & research
  • LLM2×
  • Machine learning
Security & compliance
  • Security2×
Cloud & infra
  • Distributed systems

Words the posting leans on

  • data17×
  • systems10×
  • highly9×
  • massive8×
  • simulation8×
  • complex6×
  • architecture5×
  • experience5×
  • processing5×
  • batch4×
  • context4×
  • customer4×
  • models4×
  • build3×
  • clearance3×
  • data infrastructure3×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Scale AI

every open role at Scale AI

How this page was made

An automated read of a public job posting, fetched August 24, 2026 and last changed by Scale AI on August 4, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.