Staff/Sr. ML Infrastructure Engineer, Foundation Model Compute Infra
Core
Design and build large-scale infrastructure for foundation model training, fine-tuning, evaluation, and inference.
Role type
Staff/Senior IC ML Infrastructure Engineer
Builds
Large-scale model serving and fine-tuning services, accelerator integration, and compute infrastructure for Apple's AI workloads
Domain
AI/ML Infrastructure, Distributed Systems, Cloud Computing
Deliverable
infrastructure
Required skills
Distributed systems design, Cloud infrastructure, Python, Go, C++, Accelerator hardware (TPU/GPU), Kubernetes, Container orchestration
Preferred skills
Scheduler/resource manager design, Frameworks (JAX, PyTorch, TensorFlow, Ray, Pathways, vLLM), Multi-tenant cloud operations, Performance optimization, Fault tolerance
Technologies
Kubernetes, TPU, GPU, Ray, Pathways, Beam, vLLM
Responsibilities
Design and evolve model serving and fine-tuning services; Develop infrastructure for deployment, autoscaling, and traffic management; Optimize latency, throughput, and accelerator utilization; Onboard and benchmark new accelerator technologies; Collaborate on integrating technologies like Pathways and Ray; Mentor engineers and influence technical direction
Seniority
Staff/Senior, hands-on IC with mentorship