CareerPlanSign in

ML Framework (MetalLM) Engineer, Graphics, Game and ML

Cupertino, United States of America💼 Full-time🗓 2026-03-27 → 2026-09-28

Core

Building scalable, efficient, production-grade ML inference frameworks for GenAI applications (LLMs) on Apple Silicon server hardware.

Role type

Senior IC ML Framework Engineer (GPU/Graphics)

Builds

Custom-built server ML inference frameworks for Private Cloud Compute

Domain

Server ML, GPU programming, Apple Silicon

Deliverable

production ML models

Required skills

C/C++/ObjC, GPU kernel development, distributed inference techniques, system-level programming, computer architecture

Preferred skills

Graph compilers (CuTE, CuTile, Triton, OpenXLA, LLVM), LLM/Diffusion model architectures

Responsibilities

Optimize code for efficient/scalable ML inference using distributed compute strategies; Develop kernel and compiler-level optimizations; Apply model optimization techniques (speculation, quantization, compression); Collaborate with hardware/compiler/systems teams; Analyze and improve performance metrics (latency, memory footprint, compute efficiency)

Seniority

Senior, hands-on IC

Sourced via apple · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.