Site Reliability Engineer (SRE) - Early Talent
Core
Support day-to-day SRE operations and execute well-defined projects for Nebius's full-stack AI cloud platform.
Role type
Junior Site Reliability Engineer (Early Talent)
Builds
AI cloud infrastructure components including GPU orchestration, inference optimization, compute, storage, and networking.
Domain
Cloud Infrastructure / AI / GPU Orchestration
Deliverable
production ML models | infrastructure
Required skills
Linux fundamentals (networking, cgroups, eBPF), Python/Go/C++/Java programming, Git, Kubernetes, Terraform, Helm, networking basics (Ethernet, IP, TCP/UDP, routing)
Responsibilities
Assist in day-to-day SRE operations tasks, execute tasks from the backlog, create tests for changes, write technical documentation, track Jira tasks, study system fundamentals, learn internal tools and workflows.
Seniority
Junior, Early Career