Machine Learning Compute Efficiency Lead, Infrastructure & Planning
Core
Lead compute efficiency strategy for Apple's large-scale ML inference workloads across GPUs, TPUs, and custom silicon to reduce costs and maximize performance.
Role type
Senior Architect, ML Infrastructure & Compute Efficiency
Builds
Scalable, cost-effective inference platform for Apple Intelligence and foundation models
Domain
Cloud Infrastructure, Machine Learning Systems, Hardware Optimization
Deliverable
production ML models
Required skills
ML infrastructure architecture, GPU/TPU utilization optimization, cluster scheduling, capacity planning, distributed training concepts, root cause analysis, cross-org technical leadership
Preferred skills
Foundation model serving experience, FinOps, TCO modeling, Kubernetes/Slurm, PyTorch/JAX, technical negotiation
Technologies
GPU, TPU, Apple Silicon, Kubernetes, Slurm, PyTorch, JAX
Responsibilities
Own ML compute management for inference workloads; Develop resource strategies with ML engineering teams; Optimize workloads for performance and cost reduction; Architect solutions for capacity allocation and scheduling; Advocate for ML platform requirements to infrastructure providers.
Seniority
Senior, hands-on IC with strategic influence