Platform Site Reliability Engineer
Core
Design and operate AI-native cloud platforms, managing GPU-native workloads, multi-tenant control planes, and high-performance AI systems.
Role type
Senior Platform Site Reliability Engineer (Kubernetes & Linux)
Builds
AI infrastructure platform, Kubernetes clusters, observability stack, and automation tooling.
Domain
Cloud Infrastructure, AI Systems, High-Performance Computing
Deliverable
production ML models | infrastructure
Required skills
Kubernetes administration, Linux system tuning, infrastructure as code, observability stack management, incident management, networking fundamentals, automation scripting
Preferred skills
AI workload orchestration, HPC supportability, ITSM frameworks, mentorship
Technologies
Kubernetes, Prometheus, Grafana, Bash, Python, Ansible, Linux (Ubuntu), TCP/IP, DNS, DHCP
Responsibilities
Deploy and manage scaled Kubernetes clusters for AI workloads; optimize Linux system configuration and kernel parameters; build automation scripts and IaC for platform lifecycle; maintain observability stack and conduct incident postmortems; mentor junior engineers and consult on operational requirements.
Seniority
Senior, hands-on IC with mentorship responsibilities