Senior Site Reliability Engineer (SRE)
Core
Ensure fault-tolerance, scale, and uninterrupted operations for Nebius's full-stack AI cloud platform using cutting-edge cloud technology.
Role type
Senior Site Reliability Engineer (SRE)
Builds
AI cloud platform infrastructure supporting data and model training through production deployment
Domain
Cloud infrastructure, AI/ML compute, distributed systems
Deliverable
production ML models | infrastructure
Required skills
Go, Python, C++, Unix systems, network technology, containerization, configuration management, CI/CD processes
Preferred skills
backend development, high-load distributed systems design, multi-cloud platforms
Technologies
Ansible, Salt, Terraform, Docker, K8s, Helm
Responsibilities
Ensure fault-tolerance, scale, and uninterrupted operations for the service; Use cutting-edge cloud technology to solve infrastructure problems; Implement and improve CI/CD processes
Seniority
Senior, hands-on IC