Manager SRE
Core
Leading the architecture, design, and rollout of an internal developer platform spanning CI/CD, tooling, and infrastructure as code while managing a team of SWEs and SREs to ensure reliability and performance.
Role type
Manager, Site Reliability Engineering (SRE)
Builds
Internal developer platform (IDP), automated fleet management features, and production infrastructure
Domain
Cloud-native infrastructure, DevOps, Identity Security
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud native environment management, CI/CD orchestration, Kubernetes operations, Infrastructure as Code (IaC), Linux fundamentals, Networking concepts, Security background, Team leadership, Project management
Preferred skills
Java/Tomcat experience, Containerized services management, AWS expertise (EC2, ECS, KMS, Kinesis, RDS), Vulnerability scanning, Cloud spend optimization
Technologies
Kubernetes, AWS (EC2, ECS, KMS, Kinesis, RDS), CI/CD tools, Linux, Java, Tomcat
Responsibilities
Leading architecture and rollout of internal developer platform, Enhancing automation for fleet management, Triaging and troubleshooting complex production issues, Mentoring and managing a team of SWEs and SREs, Partnering with stakeholders on reliability and delivery velocity metrics, Supporting 24x7 on-call rotation
Seniority
Manager, hands-on IC with team leadership