Software Engineer, ML Platform
Core
Build infrastructure that turns real product usage into better models and keeps research moving fast on large GPU fleets.
Role type
Senior IC ML Platform Engineer (distributed systems & infrastructure)
Builds
Core platform systems for ML researchers and product engineers, including telemetry pipelines, data platforms, observability tools, and GPU cluster scheduling.
Domain
AI/ML infrastructure, distributed systems, cloud computing
Deliverable
production ML models | infrastructure
Required skills
distributed systems, infrastructure software engineering, Linux, cloud and/or bare metal, Kubernetes, Ray, data pipelines, scheduling/orchestration
Preferred skills
event ingestion, product analytics pipelines, OpenTelemetry, Spark, Flink, GPU cluster scheduling, job queues, experiment monitoring
Technologies
Kubernetes, Ray, Spark, Flink, OpenTelemetry, Linux
Responsibilities
Design, build, and operate core platform systems used daily by ML researchers; Partner with research to turn recurring pain into durable infrastructure; Own reliability, performance, and developer experience for systems; Ship iteratively in a flat, high-ownership environment.
Seniority
Senior, hands-on IC