CareerPlanGet AI match score →

AI Platform Support Engineer (US)

San Francisco, California💼 Full-time💰 $115,000–$115,000🗓 2026-06-11 → 2026-07-31

Core

Technical partner supporting ML engineers running large-scale training and inference workloads on cloud infrastructure, Kubernetes, and GPU platforms.

Role type

Senior IC AI Platform Support Engineer

Builds

Production ML training and inference systems for solo researchers, startups, and large enterprises

Domain

AI Infrastructure / Cloud Computing / Distributed Systems

Deliverable

infrastructure

Required skills

Kubernetes, Linux systems, distributed systems, observability tools (Prometheus/Grafana/OpenTelemetry, debugging ML infrastructure (PyTorch, CUDA, NCCL), GPU orchestration, performance tuning, root cause analysis

Preferred skills

Large-scale model training, distributed scheduling (Ray/Kubeflow/Slurm), InfiniBand/RDMA, bare metal infrastructure, storage systems, Python automation

Technologies

Kubernetes, PyTorch, CUDA, NCCL, Prometheus, Grafana, OpenTelemetry, Ray, Kubeflow, Slurm, InfiniBand, RDMA

Responsibilities

Partner with customer engineering teams to diagnose and resolve complex distributed systems and ML infrastructure issues; Act as technical advisor during high impact incidents; Investigate failures involving distributed training, GPU allocation, networking, and storage; Build internal tooling, automation, and documentation to improve reliability; Contribute to post-incident reviews and operational improvements

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗