Platform Engineer
Core
Designing and maintaining robust, observable, secure, and scalable infrastructure systems to ensure the reliability and reproducibility of ML workloads.
Role type
Senior Platform Engineer (ML Infrastructure)
Builds
Resilient build and deployment systems, observability stacks, and secure release processes for research and production ML environments.
Domain
Artificial Intelligence / Machine Learning Infrastructure
Deliverable
infrastructure
Required skills
High-performance compute environment management, Infrastructure as Code (Terraform, Ansible), Container orchestration (Kubernetes, Slurm), CI/CD pipeline design, Incident response and root-cause analysis, Chaos engineering, Load testing, Compliance and audit standards.
Preferred skills
Software release engineering for ML/AI systems, GPU management and workload optimization, Backend development for ML model serving (vLLM, Ray, SGLang, Triton).
Technologies
AWS, GCP, Terraform, Ansible, Docker, Apptainer, Kubernetes, Slurm, vLLM, Ray, SGLang, Triton
Responsibilities
Building and improving observability systems (monitoring, logging, alerting), Managing Infrastructure as a Service and CI/CD, Designing resilient build and deployment systems, Implementing secure release processes with auditability, Leading incident response and postmortems, Collaborating with ML engineers and DevOps teams.
Seniority
Senior, hands-on IC