AI Platform Support Engineer (US)
Core
Technical partner supporting ML engineers running large-scale training and inference workloads on cloud infrastructure, Kubernetes, and GPU platforms.
Role type
Senior IC AI Platform Support Engineer
Builds
Production ML training and inference systems for solo researchers, startups, and large enterprises
Domain
AI Infrastructure / Cloud Computing / Distributed Systems
Deliverable
infrastructure
Required skills
Kubernetes, Linux systems, distributed systems, observability tools (Prometheus/Grafana/OpenTelemetry, debugging ML infrastructure (PyTorch, CUDA, NCCL), GPU orchestration, performance tuning, root cause analysis
Preferred skills
Large-scale model training, distributed scheduling (Ray/Kubeflow/Slurm), InfiniBand/RDMA, bare metal infrastructure, storage systems, Python automation
Technologies
Kubernetes, PyTorch, CUDA, NCCL, Prometheus, Grafana, OpenTelemetry, Ray, Kubeflow, Slurm, InfiniBand, RDMA
Responsibilities
Partner with customer engineering teams to diagnose and resolve complex distributed systems and ML infrastructure issues; Act as technical advisor during high impact incidents; Investigate failures involving distributed training, GPU allocation, networking, and storage; Build internal tooling, automation, and documentation to improve reliability; Contribute to post-incident reviews and operational improvements
Seniority
Senior, hands-on IC