Sr. IT Site Reliability Software Engineer
Databricks · Costa Rica · Infrastructure · listed May 19, 2026
The shape of it
Seniority
Senior
Experience asked
5+ years
Where
Not stated
Requirements listed
7
Length
658 words
In the posting’s own words
As a Site Reliability Engineer (SRE) , you will bridge the gap between software engineering and systems architecture. You will be a core contributor to the IT Infrastructure team, owning the evolution of core infrastructure and observability platforms. This role requires a strong software engineering mindset and deep technical breadth to deliver high-quality, scalable solutions for "immature" system problems. Your focus will be on building resilient, automated infrastructure that empowers development teams and ensures our cloud environment is cost-optimized, secure, and highly available.
What it asks for · 7
- Software Engineering Expertise: 5+ years of production-level experience with strong proficiency in Python (non-negotiable).
- Infrastructure as Code (IaC): Expert-level proficiency in Terraform (modules, state management) or Pulumi.
- Cloud & Containers: Hands-on experience with AWS, Azure, or GCP, along with Kubernetes, Docker, and containerization concepts.
- Observability Mindset: Deep understanding of observability pillars (logging, metrics, tracing) and experience with tools such as Datadog , Prometheus, or ELK.
- Distributed Systems: Proficiency in running systems using concepts like Kafka or messaging queues.
- CI/CD Proficiency: Advanced knowledge of GitHub Actions and GitHub Runners.
- Independent Execution: Ability to take ownership of ambiguous projects, follow a vision set by tech leads, and execute independently with minimal guidance.
What the job covers
- Architect and Automate: Design and deploy production-grade infrastructure on cloud platforms (AWS/Azure) using Infrastructure as Code (IaC) tools like Terraform or Pulumi.
- Reliability and Performance Engineering: Optimize system performance, architecture, and scaling to ensure maximum uptime and minimal latency for critical IT services.
- CI/CD Excellence: Architect robust deployment pipelines (e.g., GitHub Actions), managing both hosted and self-hosted runners for specialized build requirements.
- Observable by Default: Create underlying infrastructure to ensure new internal applications are secure and have logging, metrics and alerts enabled by default.
- Agentic ToolingI: Build internal AI plugins, and automation scripts to streamline developer workflows and enhance operational efficiency.
- Incident Response: Focus on subsequent data usage, incident management workflows, and creating necessary dashboards to maintain service health. Participate in a shared on-call rotation, leading rapid incident response and technical troubleshooting for production outages.Facilitate blameless post-mortems to identify root causes and implement permanent preventive engineering solutions.
- Partner Cross-Functionally: Collaborate with Security, Engineering, and Support teams to deliver real business outcomes.
Tools and skills named
Cloud & infra
- Observability3×
- AWS2×
- Azure2×
- CI/CD2×
- GitHub Actions2×
- Site reliability2×
- Terraform2×
- Datadog
- Distributed systems
- Docker
- GCP
- Kafka
- Kubernetes
- Prometheus
Data
- Databricks
Languages
- Python
Security & compliance
- Security
Ways of working
- On-call
Words the posting leans on
- infrastructure8×
- engineering6×
- systems5×
- proficiency4×
- cloud3×
- ensure3×
- experience3×
- observability3×
- services3×
- software engineering3×
- applications2×
- architect2×
- architecture2×
- build2×
- ci/cd2×
- code iac2×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.