Senior Site Reliability Engineer
Core
Senior SRE ensuring security, reliability, scalability, and operational excellence of cloud platforms and microservices.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud-based microservices on Multi-Cloud Platform (AWS, OCI, Azure, GCP)
Domain
Networking, IoT, Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes, Microservices, Cloud Operations, Disaster Recovery, Observability, Python, Go, Bash, Java, SLA/SLO/SLI definition, Incident Response, Technical Documentation, Security Compliance (ISO27001, SOC2, GDPR)
Preferred skills
Expert-level Cloud Certifications (AWS/Azure/GCP Solutions Architect), Container Orchestration
Technologies
Kubernetes, AWS, OCI, Azure, GCP, Python, Go, Bash, Java
Responsibilities
Implement and operate Microservices on Kubernetes; Deploy services to Multi-Cloud Platform; Perform Load Tests and Chaos Tests; Build Observability; Write and Execute Disaster recovery plans; Analyze and resolve production risks; Write and maintain automation scripts; Define and maintain KPIs (SLA/SLO/SLI); Create and maintain technical documentation; Guarantee adherence to security and compliance standards; Lead incident response efforts; Perform post-incident analysis; Assist with product/technology selection; Mentor less senior team members; Participate in On-call rotation.
Seniority
Senior, hands-on IC