Principal ML Platform Engineer
Core
Designing and operating the infrastructure and tooling that enables researchers and product teams to train, serve, and deploy generative models reliably and efficiently.
Role type
Principal ML Platform Engineer (Systems & Infrastructure)
Builds
Research infrastructure, production serving systems, internal tooling, and platform interfaces for generative models
Domain
AI/ML infrastructure, cloud systems, distributed computing
Deliverable
infrastructure
Required skills
Production system design, systems mindset, cloud infrastructure, Linux, infrastructure automation, Kubernetes, distributed workloads, Python, observability, debugging, architectural tradeoffs
Preferred skills
ML infrastructure operation, GPU-based systems, workflow orchestration, agentic systems, performance optimization
Technologies
Kubernetes, Terraform, Datadog, GitHub Actions, Temporal, Python
Responsibilities
Design platform systems for model training and serving; build infrastructure for reliable and cost-efficient ML workloads; develop internal tools for human and agent operation; architect model deployment and serving; improve scheduling, monitoring, and debugging of GPU workloads; drive observability and automation improvements; collaborate with researchers to solve pain points; contribute to technical direction.
Seniority
Principal, hands-on IC with significant ownership