Senior Software Engineer, AI Operations, GPS
Scale AI · Doha, Qatar · GPS Engineering · listed August 20, 2026
The shape of it
Seniority
Senior
Where
Not stated
Requirements listed
3
Length
864 words
In the posting’s own words
As a Senior Software Engineer, AI Ops at Scale AI, you will own the long-term technical health, performance, and stability of AI solutions deployed across our strategic public sector partners.
What it asks for · 3
- Technical Stack: Advanced proficiency in Python, SQL, REST/gRPC APIs, and cloud architecture (AWS, Azure, or GCP). Hands-on experience with MLOps tooling, vector databases, and LLM orchestration frameworks (e.g., LangChain, LlamaIndex).
- Engineering Mindset: A drive to build systematic, automated fixes rather than applying temporary patches. Strong grasp of CI/CD for machine learning pipelines.
- Client Acumen & Boundary Control: Strong technical communication skills with the ability to manage client expectations, defend operational boundaries (Maintenance vs. Evolution), and advise on long-term system roadmaps.
What the job covers
- Handover Gate & Onboarding: Act as the technical gatekeeper during the formal transition from Delivery to Maintenance. Conduct deep-dive reviews to ensure baseline code, prompts, and architecture meet strict maintainability and documentation standards before sign-off.
- Tiered SLA & Incident Management: Own technical response and resolution targets across multi-tiered service models (from Business-Hours Essential to 24/7 Mission-Critical). Lead Incident Governance, Root Cause Analysis (RCA), and P1/P2 mitigations within strict active support windows.
- AI Lifecycle Governance: Monitor production model performance, latency, and data drift. Manage prompt configuration repositories to maintain behavioral consistency and perform regression testing when LLM providers update underlying endpoints.
- Request Classification & Technical Scope: Operationalize the boundary between Routine Maintenance (In-Scope) and System Evolution (Out-of-Scope). Assess incoming client requests and run comparative benchmarking on new AI models.
- Automation & Reliability Engineering: Eliminate operational toil by engineering self-healing data pipelines, automated RAG indexing syncs, and telemetry tooling. Influence upstream "Delivery" teams to adopt architectural patterns that simplify ongoing maintenance.
- Client Technical Interface: Serve as the senior technical point of contact for government and enterprise IT leads. Translate technical AI concepts (data drift, prompt versioning, API deprecation) into clear business impacts for non-technical stakeholders.
Tools and skills named
Cloud & infra
- AWS
- Azure
- CI/CD
- GCP
- Site reliability
Models & research
- LLM2×
- Machine learning
Ways of working
- Technical writing2×
- Testing
Frameworks
- gRPC
- REST
Languages
- Python
- SQL
Data
- Data pipelines
Words the posting leans on
- technical10×
- engineering6×
- client5×
- maintenance5×
- model5×
- governance4×
- operational4×
- prompt4×
- visa4×
- data drift3×
- delivery3×
- pipelines3×
- qatar3×
- system3×
- applying2×
- architecture2×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Scale AI
- AI Advisory ConsultantSan Francisco, CA; New York, NY
- AI Advisory PrincipalSan Francisco, CA; New York, NY
- AI Applications Ops Manager, GPSDoha, Qatar
- AI Builder InternSan Francisco, CA
- AI Infrastructure Engineer, Model Serving PlatformSan Francisco, CA; New York, NY
- AI Infrastructure Engineer, Sandbox PlatformLondon, UK