Machine Learning Engineer, Global Public Sector
Scale AI · Doha, Qatar; London, UK · GPS Engineering · listed May 1, 2024
The shape of it
Seniority
Not stated
Where
Not stated
Requirements listed
5
Length
1,035 words
In the posting’s own words
Your role is to move beyond off-the-shelf implementations. You will lead the research into Agent Design, Reliability, and AI Safety, developing novel system architectures that power high-stakes government applications. You will be the bridge between a research paper and a production-ready system that functions at scale.
What it asks for · 5
- Engineering Rigour: Exceptional proficiency in Python and experience building agentic harnesses or AI infrastructure. You write production-ready code that is modular, scalable, and reliable.
- Applied Research Mindset: A track record of taking theoretical AI concepts and turning them into functional prototypes or products. You know how to read a paper and determine if its methods are actually viable for a production system.
- Evaluation Expertise: Experience in LLM benchmarking, red-teaming, or building evaluations that go beyond standard academic datasets.
- Advanced Degree: A Master’s or PhD in Computer Science, Mathematics, or a related field (with a focus on ML) is preferred, but we value demonstrated impact and engineering excellence.
- Agentic Systems Expert: Deep experience in building multi-agent systems, including chain-of-thought optimisation and tool-calling reliability.
Also a plus
- Agentic Systems Expert: Deep experience in building multi-agent systems, including chain-of-thought optimisation and tool-calling reliability.
- Sovereign AI Experience: Experience working with highly regulated data environments, on-premise deployments, or sensitive government use cases.
- Inference Optimisation: Knowledge of how to optimise model performance for environments with limited GPU capacity or specific latency requirements.
- Zero-to-One Mindset: You are comfortable navigating ambiguity and enjoy defining research directions from scratch to solve a specific product or mission need.
What the job covers
- Architect Agentic Systems: Design and build agent architectures, the harnesses, tool-use protocols, and logic flows that allow LLMs to function as reliable, autonomous agents in complex workflows.
- Drive Reliability & Safety: Research and implement robust evaluation frameworks. This includes red-teaming for sovereign AI requirements and developing strategies to mitigate hallucinations in regulated data environments.
- Synthesise Deep Research: Build agents capable of autonomous information synthesis and long-horizon reasoning, enabling users to analyse massive datasets and extract actionable insights.
- Optimize for Niche Domains: Evaluate and adapt models for specialised use cases, such as LLM reasoning for low-resource languages, complex OCR tasks, or working in GPU-constrained environments
- Build Evaluation Frontiers: Create new, automated benchmarks that define what success looks like for AI in the public sector, ensuring our systems meet the highest standards of accuracy and sovereignty.
- Consult as a Technical Authority: Act as a subject matter expert for public sector leaders, advising on the practical limits, safety requirements, and performance trade-offs of emerging AI technologies.
Degree language
- Advanced Degree: A Master’s or PhD in Computer Science, Mathematics, or a related field (with a focus on ML) is preferred, but we value demonstrated impact and engineering excellence.
Tools and skills named
Models & research
- LLM4×
- GPU2×
- Inference2×
- Machine learning2×
- Evaluations
Languages
- Python
Security & compliance
- Security
Ways of working
- Technical writing
Words the posting leans on
- research10×
- systems10×
- evaluation7×
- agent5×
- experience5×
- impact4×
- llm4×
- model4×
- reliable4×
- visa4×
- agentic systems3×
- build3×
- complex3×
- design3×
- developing3×
- engineering3×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Scale AI
- AI Advisory ConsultantSan Francisco, CA; New York, NY
- AI Advisory PrincipalSan Francisco, CA; New York, NY
- AI Applications Ops Manager, GPSDoha, Qatar
- AI Builder InternSan Francisco, CA
- AI Infrastructure Engineer, Model Serving PlatformSan Francisco, CA; New York, NY
- AI Infrastructure Engineer, Sandbox PlatformSan Francisco, CA; Seattle, WA; New York, NY