CareerPlanSign in

Platform Engineer

San Francisco💼 Full-time🗓 2026-03-17 → 2026-09-25

Core

Designing and maintaining robust, observable, secure, and scalable infrastructure systems to ensure the reliability and reproducibility of ML workloads.

Role type

Senior Platform Engineer (ML Infrastructure)

Builds

Resilient build and deployment systems, observability stacks, and secure release processes for research and production ML environments.

Domain

Artificial Intelligence / Machine Learning Infrastructure

Deliverable

infrastructure

Required skills

High-performance compute environment management, Infrastructure as Code (Terraform, Ansible), Container orchestration (Kubernetes, Slurm), CI/CD pipeline design, Incident response and root-cause analysis, Chaos engineering, Load testing, Compliance and audit standards.

Preferred skills

Software release engineering for ML/AI systems, GPU management and workload optimization, Backend development for ML model serving (vLLM, Ray, SGLang, Triton).

Technologies

AWS, GCP, Terraform, Ansible, Docker, Apptainer, Kubernetes, Slurm, vLLM, Ray, SGLang, Triton

Responsibilities

Building and improving observability systems (monitoring, logging, alerting), Managing Infrastructure as a Service and CI/CD, Designing resilient build and deployment systems, Implementing secure release processes with auditability, Leading incident response and postmortems, Collaborating with ML engineers and DevOps teams.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.