ML Framework (MetalLM) Engineer
Core
Building high-performance, distributed inference and training frameworks for GenAI applications (LLMs) on Apple Silicon server hardware.
Role type
Senior IC ML Framework Engineer (GPU Systems)
Builds
Custom ML inference and training frameworks optimized for Apple Silicon data center GPUs
Domain
Cloud Compute / GPU Acceleration / Machine Learning
Deliverable
production ML models
Required skills
C/C++/ObjC, GPU kernel development, Metal runtime, distributed training/inference, system-level programming, computer architecture, model optimization (speculation, quantization, compression)
Preferred skills
Graph compilers (Triton, OpenXLA, LLVM/MLIR), AI framework contributions (PyTorch, JAX, Tensorflow), machine learning fundamentals
Technologies
Metal, CUDA, PyTorch, JAX, Triton, OpenXLA, LLVM/MLIR
Responsibilities
Optimize code for efficient and scalable ML inference using distributed compute strategies; Develop kernel and compiler level optimizations; Apply advanced model optimization techniques; Collaborate with hardware, compiler, and systems teams; Analyze and improve performance metrics (latency, memory footprint, compute efficiency); Implement features of Metal device backend for ML training acceleration
Seniority
Senior, hands-on IC
