Research Engineer, Takeoff Intel

Anthropic · Remote-Friendly (Travel Required) | San Francisco, CA · AI Research & Engineering · listed September 8, 2026

The shape of it

Seniority
Not stated
Where
Hybrid
Stated pay
$350,000 – $850,000 USD
Requirements listed
6
Length
1,385 words

In the posting’s own words

Our work appears in Anthropic's model system cards (we own the AI R&D capability assessments and adapted Epoch's Capabilities Index to our evals); all the data in When AI Builds Itself comes from our team. Internally, our measurements shape research priorities and safety planning; externally, they contribute to Anthropic's public reporting on the pace of AI progress and to collaborations with third-party evaluators. We're a small team that works closely with pretraining, RL, economics, and policy researchers across the company. If you're passionate about measurement accuracy, and feel urgency about safety and situational awareness, you should consider joining us.

What it asks for · 6

  • Have shipped an evaluation, data product, or research library end to end
  • Prototype fast and are comfortable throwing code away
  • Handle messy, large-volume data without over-engineering
  • Have run experiments on large language models, not just moved their outputs around
  • Can work from a vague question rather than a spec
  • Communicate results clearly and collaborate closely with the researchers whose questions your instruments answer

Also a plus

  • Built evaluation harnesses or benchmark infrastructure for LLMs
  • Experience with large-scale ML or data infrastructure (self-driving, observability, or similar) alongside ML exposure
  • Built tools or libraries that other researchers rely on
  • A track record of catching what AI-written code gets wrong

What the job covers

  • Design, build, and run capability evaluations and measurement instruments at scale
  • Build the data and analysis pipelines that turn large volumes of model outputs and telemetry into reliable metrics
  • Prototype new instruments fast, validate them, and decide what to keep
  • Review and supervise AI-written code as a normal part of the workflow
  • Work closely with research scientists on the team and with partner teams to define what's worth measuring
  • Contribute to internal write-ups and public reporting

Tools and skills named

Models & research
  • Evaluations5×
  • LLM4×
  • Machine learning4×
Cloud & infra
  • Observability2×

Words the posting leans on

  • data12×
  • build10×
  • instruments9×
  • capability8×
  • evaluation8×
  • research8×
  • model6×
  • question6×
  • closely5×
  • infrastructure5×
  • prototype5×
  • researchers5×
  • run5×
  • system cards5×
  • ai-written code4×
  • built4×

Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.

The posting, your resume, and the gaps between them. One click loads all three.

More open at Anthropic

every open role at Anthropic

How this page was made

An automated read of a public job posting, fetched September 8, 2026 and last changed by Anthropic on September 8, 2026. Every list above is pulled from the posting’s own sentences — nothing rewritten, nothing added, no judgment about the role or the company. Counts and seniority are read off the text by rule, so they can be wrong where the posting is unusual. The original is the only thing that binds. Openings close without warning; check the source before spending an evening on it.