CareerPlanSign in

Machine Learning Compute Efficiency Lead, Infrastructure & Planning

Cupertino, United States of America💼 Full-time🗓 2026-04-23 → 2026-09-28

Core

Lead compute efficiency strategy for Apple's large-scale ML inference workloads across GPUs, TPUs, and custom silicon to reduce costs and maximize performance.

Role type

Senior Architect, ML Infrastructure & Compute Efficiency

Builds

Scalable, cost-effective inference platform for Apple Intelligence and foundation models

Domain

Cloud Infrastructure, Machine Learning Systems, Hardware Optimization

Deliverable

production ML models

Required skills

ML infrastructure architecture, GPU/TPU utilization optimization, cluster scheduling, capacity planning, distributed training concepts, root cause analysis, cross-org technical leadership

Preferred skills

Foundation model serving experience, FinOps, TCO modeling, Kubernetes/Slurm, PyTorch/JAX, technical negotiation

Technologies

GPU, TPU, Apple Silicon, Kubernetes, Slurm, PyTorch, JAX

Responsibilities

Own ML compute management for inference workloads; Develop resource strategies with ML engineering teams; Optimize workloads for performance and cost reduction; Architect solutions for capacity allocation and scheduling; Advocate for ML platform requirements to infrastructure providers.

Seniority

Senior, hands-on IC with strategic influence

Sourced via apple · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.