AI Platform Engineer
Core
Build the bridge between advanced AI infrastructure and teams developing, training, and deploying models and AI services by creating a secure, automated, and attractive developer experience on top of GPU clusters.
Role type
Senior IC AI Platform Engineer
Builds
Standardized Kubernetes-based platform services for AI/ML workloads, including self-service APIs, pipelines, and guardrails.
Domain
Cloud infrastructure + AI/ML platform engineering
Deliverable
production ML models | infrastructure
Required skills
Kubernetes, container platforms, DevOps, SRE, GitOps, CI/CD, Infrastructure as Code, API design, observability, policy enforcement, distributed systems, automation, GPU resource management, workload scheduling, model serving, data integration, identity and secrets management
Preferred skills
Nvidia NIM, Nvidia AI Enterprise, GPU operators, GPU-aware scheduling, multi-tenant Kubernetes, resource isolation, model registries, vector databases, FinOps, confidential computing, sovereign cloud
Technologies
Kubernetes, GPU clusters, MLOps, API, observability tools, CI/CD pipelines
Responsibilities
Design and evolve Kubernetes-based platform services for AI/ML workloads; create standardized workflows for development, training, experimentation, model management, inference, and lifecycle; integrate GPU resources, scheduling, storage, identity, secrets, networking, and observability; develop self-service APIs, templates, pipelines, and guardrails; automate deployment, configuration, upgrades, and policy application; collaborate with AI Infrastructure, Security, and user teams to translate workload requirements into platform capabilities; monitor stability, resource utilization, and user experience to drive continuous improvements.
Seniority
Senior, hands-on IC