CareerPlanGet AI match score →

Site Reliability Engineer

Costa Rica💼 Full-time🗓 2026-05-19 → 2026-07-31

Core

Bridge software engineering and systems architecture to build resilient, automated infrastructure and observability platforms for IT services.

Role type

Site Reliability Engineer (SRE)

Builds

Production-grade infrastructure, CI/CD pipelines, observability tooling, and internal AI automation scripts.

Domain

Cloud Infrastructure & Observability

Deliverable

production ML models | infrastructure

Required skills

Python, Terraform, Kubernetes, AWS/Azure/GCP, Datadog, Prometheus, Kafka, GitHub Actions

Preferred skills

Pulumi, ELK stack, distributed systems concepts

Technologies

Terraform, Pulumi, AWS, Azure, GCP, Kubernetes, Docker, Datadog, Prometheus, ELK, Kafka, GitHub Actions

Responsibilities

Design and deploy infrastructure using IaC; optimize system performance and scaling; architect CI/CD pipelines; create observability infrastructure; build internal AI plugins and automation scripts; manage incident response and post-mortems.

Seniority

Mid-Senior, hands-on IC

Rewrite
## Responsibilities - Architect and Automate: Design and deploy production-grade infrastructure on cloud platforms (AWS/Azure) using Infrastructure as Code (IaC) tools like Terraform or Pulumi. - Reliability and Performance Engineering: Optimize system performance, architecture, and scaling to ensure maximum uptime and minimal latency for critical IT services. - CI/CD Excellence: Architect robust deployment pipelines (e.g., GitHub Actions), managing both hosted and self-hosted runners for specialized build requirements. - Observable by Default: Create underlying infrastructure to ensure new internal applications are secure and have logging, metrics and alerts enabled by default. - Agentic Tooling: Build internal AI plugins, and automation scripts to streamline developer workflows and enhance operational efficiency. - Incident Response: Focus on subsequent data usage, incident management workflows, and creating necessary dashboards to maintain service health. Participate in a shared on-call rotation, leading rapid incident response and technical troubleshooting for production outages. Facilitate blameless post-mortems to identify root causes and implement permanent preventive engineering solutions. - Partner Cross-Functionally: Collaborate with Security, Engineering, and Support teams to deliver real business outcomes. ## Requirements - Software Engineering Expertise: 5+ years of production-level experience with strong proficiency in Python (non-negotiable). - Infrastructure as Code (IaC): Expert-level proficiency in Terraform (modules, state management) or Pulumi. - Cloud & Containers: Hands-on experience with AWS, Azure, or GCP, along with Kubernetes, Docker, and containerization concepts. - Observability Mindset: Deep understanding of observability pillars (logging, metrics, tracing) and experience with tools such as Datadog, Prometheus, or ELK. - Distributed Systems: Proficiency in running systems using concepts like Kafka or messaging queues. - CI/CD Proficiency: Advanced knowledge of GitHub Actions and GitHub Runners. - Independent Execution: Ability to take ownership of ambiguous projects, follow a vision set by tech leads, and execute independently with minimal guidance. ## Nice to Have - None specified. ## Benefits - About Databricks: Databricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook.
Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗