Lead/Manager Site Reliability Engineering Team (Amsterdam)
Core
Lead a team of Site Reliability Engineers to keep user-facing services and production systems running smoothly for Together AI's massive concurrent user base.
Role type
Senior IC/Lead Site Reliability Engineer
Builds
Scalable infrastructure and monitoring systems for AI research and deployment
Domain
Artificial Intelligence Infrastructure / Cloud Systems
Deliverable
infrastructure
Required skills
SRE operations, team leadership, Ansible, Terraform, Kubernetes, distributed systems, observability, cloud services, incident management, system design
Preferred skills
Algorithms, production debugging, architecture improvement
Technologies
Ansible, Terraform, Kubernetes, PagerDuty
Responsibilities
Manage and coach the SRE team, build and run infrastructure, design operational processes, debug production issues, plan infrastructure growth
Seniority
Senior, hands-on IC with leadership responsibilities