ML Infrastructure Engineer - ML Compute Capacity
Core
Design, build, and operate production systems to optimally distribute compute resources across Apple's largest accelerator fleet for ML training and inference.
Role type
Senior IC ML Infrastructure Engineer (Compute Capacity)
Builds
Production systems for demand/capacity planning, telemetry, observability, and self-service platforms for fleet management.
Domain
High-performance computing, distributed systems, ML infrastructure
Deliverable
production ML models | infrastructure
Required skills
Machine learning infrastructure on GPUs/TPUs, Python/Go, data pipelines, large-scale data querying, observability tools, problem-framing, CS fundamentals
Preferred skills
Kubernetes at production scale, modern web frameworks, accelerator utilization patterns, capacity planning, cost attribution, FinOps systems
Technologies
Trino, PostgreSQL, Elasticsearch, Prometheus, Grafana, Kubernetes, React
Responsibilities
Build and operate demand and capacity planning systems; Build data pipelines and telemetry systems for fleet-wide utilization and cost data; Develop observability infrastructure for real-time fleet health signals; Drive innovation in forecasting and supply chain management tooling; Build end-to-end tooling for actionable insights; Build self-service platforms with defined schema contracts and APIs; Engage cross-functionally with finance, supply chain, and operations teams.
Seniority
Senior, hands-on IC