Sr. Site Reliability Engineer I
Core
Building robust, scalable, and secure Kubernetes platforms and tools to support reliable engineering operations for mission-critical cloud-native services.
Role type
Senior Site Reliability Engineer (Platform Engineering)
Builds
Cloud-native Kubernetes platforms, infrastructure automation tools, and observability solutions for engineering teams.
Domain
Cloud-native infrastructure, Kubernetes, Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (AKS/EKS), Cloud platforms (Azure/AWS), Python/Go/Java/C#, Infrastructure as Code (Terraform/Pulumi), CI/CD automation, Observability (APM/Logging/Metrics), Distributed systems debugging, Technical project management
Preferred skills
Experience with both Azure and AWS, Building clustering solutions at scale, Designing tooling for SaaS/PaaS operational management
Technologies
Kubernetes, AKS, EKS, Azure, AWS, Python, Go, C#, Java, Terraform, Pulumi, APM, CI/CD platforms
Responsibilities
Build robust Kubernetes platforms and tools for rapid service provisioning; Exemplify cloud-native site reliability best practices; Write performant and maintainable code; Debug problems in cloud-native distributed systems; Scope, plan, and define technical projects; Influence engineering teams to adopt improved architectural patterns; Provide robust documentation for self-service; Continuously improve platform reliability, operability, and cost efficiency
Seniority
Senior, hands-on IC
