Staff Software Engineer, Machine Learning Platform
Stripe · Toronto · 8212 ML Foundations · listed March 12, 2026
The shape of it
Seniority
Staff
Experience asked
10+ years
Where
Not stated
Requirements listed
8
Length
879 words
In the posting’s own words
The ML Platform team builds the platforms and services that enable ML engineers and data scientists across Stripe to take data and build features and models from prototype to production—reliably, at low latency, and at scale. Our scope spans ML training infrastructure, model serving and deployment, feature computation and online serving, observability and monitoring, and agentic AI capabilities. We work closely with product teams, data scientists, and platform infrastructure teams to build powerful, flexible, and user-friendly systems that substantially increase ML velocity across the company.
What it asks for · 8
- 10+ years of professional software development experience, or equivalent domain expertise, with a solid background in service-oriented architecture and large-scale distributed systems.
- Track record of serving as a technical lead, with the ability to provide technical direction, lead multi-team initiatives, and mentor team members.
- Experience building and operating production ML platform in one or more areas such as model training, model serving, orchestration, or ML data systems, with requirements for performance, reliability, scalability, and cost efficiency.
- Strong product instincts and a deep understanding of the business context in which you operate.
- Strong communication skills with the ability to explain complex technical concepts to both technical and non-technical stakeholders.
- Demonstrated ability to work cross-functionally, collaborating effectively with ML engineers, data scientists, software engineers, product managers, and business stakeholders.
- The ability to thrive on a high level of autonomy and responsibility, and comfort operating in ambiguous environments.
- Hands-on experience using AI tools to accelerate how you work.
Also a plus
- Experience building large-scale ML training, serving, or data infrastructure for machine learning use cases, such as distributed training, model inference, feature stores, real-time feature computation, and model registries.
- Experience with distributed ML training systems, accelerator-backed compute, training data pipelines, experiment tracking, and model evaluation.
- Experience rapidly developing prototypes and iterating based on user feedback.
- Experience training and shipping machine learning models to production to solve critical business problems.
- Familiarity with LLMs, LLM application frameworks, and agentic AI patterns (e.g., tool use, multi-agent orchestration, retrieval-augmented generation).
- Familiarity with cloud services (e.g., AWS) and cloud-based AI and ML services (e.g., SageMaker, Bedrock, Databricks, OpenAI).
- Ability to synthesize ideas across the organization while setting a compelling technical vision.
- Comfortable working with geographically distributed teams.
What the job covers
- Take ownership of end-to-end architecture and system design for large, complex projects across ML Platform.
- Define technical direction for highly ambiguous projects, transforming complex user needs into long-lasting platform strategy.
- Design system architectures for the most challenging ML Platform problems in one or more areas, including AI and ML workflow orchestration, scalable CPU and GPU compute infrastructure, model training, LLM fine-tuning, low-latency model inference, large-scale feature stores, real-time monitoring, and LLM and agent orchestration.
- Turn high-leverage ideas into tangible, robust solutions that shape platform and product roadmap, combining technical excellence with creative problem-solving.
- Scope and lead large projects with significant business impact, driving them from requirements through design, implementation, and production operation.
- Work with ML engineers, data scientists, and product teams directly to translate their needs into functional requirements and scalable technical solutions.
- Arbitrate critical decisions that balance competing priorities while meeting latency, reliability, cost, and security constraints.
- Serve as a key engineering representative, engaging senior leaders across Stripe and advising the leadership team on key technical considerations related to the end-to-end ML lifecycle.
- Drive cross-team technical initiatives that improve ML development velocity and MLOps maturity across the company.
- Mentor and grow other engineers. Serve as a role model for designing, implementing, and operating great software systems.
Tools and skills named
Models & research
- Machine learning25×
- LLM4×
- Inference2×
- Fine-tuning
- GPU
Cloud & infra
- AWS
- Distributed systems
- Observability
Data
- Data pipelines
- Databricks
Product & design
- Design systems
- Roadmap
Security & compliance
- Security
Words the posting leans on
- technical14×
- model11×
- platform11×
- data10×
- product9×
- systems8×
- training8×
- experience7×
- requirements7×
- engineers6×
- feature5×
- infrastructure5×
- lead5×
- serving5×
- business4×
- data scientists4×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Stripe
- Account Executive, AI Startups (Hunter)San Francisco
- Account Executive, BridgeSF, NYC, SEA, CHI
- Account Executive, Bridge (Stripe Crypto and Stablecoins)London
- Account Executive, Commercial (Grower)Chicago
- Account Executive, Commercial HunterSingapore
- Account Executive, Commercial Hunter (Japanese Fluency)Japan