Senior Site Reliability Engineer
Core
Build and operate cloud infrastructure, delivery systems, and operational practices to enable reliable software shipping for product applications, data systems, and ML/AI workloads.
Role type
Senior Site Reliability Engineer (AI-native infrastructure)
Builds
Cloud infrastructure, CI/CD pipelines, observability platforms, and operational practices for production workloads
Domain
Pharma / AI / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS, Snowflake, Kubernetes, Docker, Terraform/OpenTofu, Python, virtual networking, incident response, root cause analysis, SLOs, automation, agentic coding systems validation
Preferred skills
Azure, GCP, Vercel, Terragrunt, MLOps, model serving, workflow orchestration, regulated environment operations
Technologies
AWS, Snowflake, Kubernetes, Docker, GitHub, Terraform, OpenTofu, Terragrunt, Python
Responsibilities
Own infrastructure and operational platform for shared engineering workloads; Build and operate secure, observable, reliable infrastructure for product apps and ML pipelines; Research and maintain core AWS infrastructure and cloud outposts; Create and optimize IaC, CI/CD pipelines, and platform patterns; Establish operational practices including SLOs, monitoring, alerting, and incident response; Partner with engineering teams to implement architecture for software and ML workloads; Use AI tools to accelerate development and validate output; Mentor engineers on SRE fundamentals
Seniority
Senior, hands-on IC
