AI Platform Support Engineer (EMEA)
Core
Technical partner to ML engineers diagnosing failures and improving reliability in large-scale distributed AI training and inference workloads.
Role type
Senior IC AI Platform Support Engineer
Builds
Production AI systems for solo researchers, startups, and large enterprises
Domain
AI Infrastructure / Cloud Computing / Distributed Systems
Deliverable
infrastructure
Required skills
Kubernetes, Linux systems, distributed systems, observability tools (Prometheus/Grafana/OpenTelemetry), PyTorch, CUDA, NCCL, GPU orchestration, networking, storage systems
Preferred skills
Ray, Kubeflow, Slurm, InfiniBand, RDMA, bare metal infrastructure, Python scripting
Responsibilities
Partner with customer engineering teams on production workloads, diagnose complex distributed systems and ML infrastructure issues, act as technical advisor during high-impact incidents, build internal tooling and automation, improve observability and troubleshooting workflows
Seniority
Senior, hands-on IC