Site Reliability Engineer (SRE)
Core
Design and implement fault-tolerant, scalable, and distributed services for research and education cloud infrastructure.
Role type
Senior Site Reliability Engineer (Infrastructure & Cloud)
Builds
Cloud infrastructure, Internal Developer Platform, and distributed services for academic and research institutions.
Domain
Research & Education / Cloud Infrastructure / Open Source
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes, Linux internals, distributed systems design, DevOps practices, containerized workloads, programming languages, OpenStack, Ceph, Ansible, GitLab CI, ArgoCD
Responsibilities
Design and implement fault-tolerant, scalable, and distributed services; handle under-the-hood investigation of legacy infrastructure and technical debt; lead projects within the team; manage services, pipelines, and tooling on the Internal Developer Platform.
Seniority
Mid-Senior, hands-on IC
