Site Reliability Engineer (SRE) AI Infrastructure (Early Career)
Core
Early-career Site Reliability Engineer supporting day-to-day operations and small projects for Nebius's full-stack AI cloud platform infrastructure.
Role type
Early Career Site Reliability Engineer (AI Infrastructure)
Builds
Full-stack AI cloud platform supporting data, model training, and production deployment
Domain
Cloud Infrastructure / AI / Networking
Deliverable
infrastructure
Required skills
Python, Go, C++, Linux Kernel, eBPF, Kubernetes, Terraform, Git, networking fundamentals (Ethernet, IP, TCP/UDP, routing)
Preferred skills
Container fundamentals, Infrastructure as Code (IaC), configuration management basics
Technologies
Kubernetes, Helm, Terraform, eBPF, DPDK, IPVS
Responsibilities
Assist in day-to-day SRE operations tasks, deploy tested and approved changes, execute small and well-defined tasks from the backlog, create tests for changes, write technical documentation, track and update Jira tasks
Seniority
Early Career / Student or Recent Graduate