Senior Site Reliability Engineer, Core AI Infrastructure
Core
Own the reliability, monitoring, and incident response lifecycle for AI infrastructure services, ensuring systems are resilient, observable, and secure at scale.
Role type
Senior Site Reliability Engineer (AI Infrastructure)
Builds
AI infrastructure services, CI/CD frameworks, and internal AI products
Domain
Cloud infrastructure, AI transformation, FinTech
Deliverable
production ML models | infrastructure
Required skills
AWS cloud infrastructure automation, Infrastructure-as-Code (Terraform, Ansible, Chef, Puppet, Salt), Container orchestration (Docker, Kubernetes), Scripting/Programming (Python, Bash, Ruby, Go), Git-based CI/CD pipelines, Incident response and root cause analysis, Generative AI usage with human oversight
Preferred skills
None stated
Technologies
AWS, Terraform, Ansible, Chef, Puppet, Salt, Docker, Kubernetes, Go, Python, Bash, Ruby, Git
Responsibilities
Own reliability, monitoring, and incident response lifecycle for AI infrastructure services; Build automation and tooling to streamline operational IT workflows and improve deployment velocity; Partner with Infrastructure and Security teams to extend CI/CD frameworks and integrate surveillance tooling; Strengthen observability and documentation standards by defining metrics and implementing monitoring solutions; Develop full-stack applications that power internal AI products and infrastructure
Seniority
Senior, hands-on IC