Staff Site Reliability Engineer, Core AI Infrastructure
Core
Own the reliability, monitoring, and incident response lifecycle for critical AI infrastructure services, ensuring systems are resilient, observable, and secure at scale.
Role type
Staff Site Reliability Engineer (AI Infrastructure)
Builds
AI infrastructure services, CI/CD frameworks, and internal AI product applications
Domain
Cloud Infrastructure / AI Operations
Deliverable
production ML models | infrastructure
Required skills
Cloud infrastructure automation, Container orchestration, Incident response, Infrastructure as Code, Scripting/Programming, CI/CD pipelines
Preferred skills
Linux administration, Log aggregation, Network security, Regulated environment experience
Technologies
AWS, Terraform, Ansible, Chef, Puppet, Salt, Docker, Kubernetes, Go, Python, Bash, Ruby, Git
Responsibilities
Own reliability, monitoring, and incident response lifecycle for AI infrastructure services; Build automation and tooling to streamline operational IT workflows; Partner with Infrastructure and Security teams to extend CI/CD frameworks; Strengthen observability and documentation standards; Develop full-stack applications for internal AI products and infrastructure
Seniority
Staff, hands-on IC