IC5 - Staff SRE
Core
Architecting and scaling the reliability, availability, and performance of critical infrastructure platforms.
Role type
Staff Site Reliability Engineer (Infrastructure)
Builds
Scalable, reliable, and secure cloud-native platform services
Domain
Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, Kubernetes, Infrastructure-as-Code, Python, Go, Bash, Cloud platforms (AWS/GCP/Azure), Observability, Capacity planning, Disaster recovery
Preferred skills
Cross-team influence, Strategic roadmap definition, Automation design, Mentorship
Technologies
Kubernetes, Python, Go, Bash, AWS, GCP, Azure
Responsibilities
Design and evolve complex infrastructure systems, Define long-term architectural vision, Lead strategic SRE initiatives, Identify systemic issues and lead root cause analysis, Champion best practices in observability, Design and implement advanced automation solutions, Mentor senior and mid-level engineers, Lead capacity planning efforts, Own and continuously improve disaster recovery strategies, Contribute to and enforce standards for infrastructure documentation
Seniority
Staff, strategic leadership with hands-on technical execution