Platform Engineer
Core
Own the platform layer between people and GPU infrastructure, managing CI/CD pipelines, Kubernetes clusters, and cloud environments for AI safety research.
Role type
Senior Platform Engineer (Infrastructure & DevOps)
Builds
CI/CD pipelines, Kubernetes clusters, internal services (artifact/container registries), and cloud environments as code.
Domain
AI Safety / Research Infrastructure / Cloud & On-premise
Deliverable
production ML models | infrastructure
Required skills
Kubernetes operations, CI/CD pipeline design, Infrastructure as Code (Terraform), Linux fundamentals, scripting (Python/Go/Bash), container security, observability (Prometheus/Grafana), IAM and network segmentation.
Preferred skills
GitOps workflows, experience with major cloud providers (AWS/GCP/Azure), on-premise infrastructure management, GPU scheduling for ML workloads, experience supporting research teams.
Technologies
Kubernetes, Terraform, Prometheus, Grafana, GitHub Actions, GitLab CI, Jenkins, Slurm, container registries.
Responsibilities
Design and run Kubernetes clusters for research workloads; define and enforce CI/CD best practices with security scanning; manage cloud environments as code; implement internal services for research teams; build observability and security controls; document and automate platform processes; participate in incident response.
Seniority
Mid-Senior, hands-on IC