Senior Site Reliability Engineer, AI Platform
Core
Design and maintain scalable, reliable cloud-native infrastructure for AI Platform services, including containerized workloads, serverless functions, and managed ML infrastructure.
Role type
Senior Site Reliability Engineer (AI Platform)
Builds
Cloud-native services, CI/CD pipelines, automation tools, and observability systems for AI Platform
Domain
Cloud Infrastructure / AI Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Python, shell scripting, Infrastructure-as-Code (CloudFormation, Terraform, ARM, SAM), CI/CD tooling, containerization (Docker, Kubernetes), cloud-native deployment patterns, observability (metrics, logs, distributed tracing), incident response, security remediation
Preferred skills
Mentoring junior engineers, designing long-term infrastructure strategy, automating toil
Technologies
Datadog, CloudWatch, Wiz, Snyk, Docker, Kubernetes, CloudFormation, Terraform, ARM, SAM
Responsibilities
Own long-term infrastructure strategy, improve operational posture of cloud-native services, instrument services with observability tooling, establish and enforce SLAs and error budgets, collaborate on security hardening, design and maintain CI/CD pipelines, automate toil, mentor junior team members
Seniority
Senior, hands-on IC