Senior Machine Learning Engineer, AI Platform
Core
Design and operate distributed systems for large-scale GPU training and low-latency inference, managing multi-tenant fleets and moving models from experiment to production.
Role type
Staff Machine Learning Platform Engineer
Builds
Shared compute and inference platform for Adobe's AI products (Firefly, Creative Cloud, Experience Cloud)
Domain
Cloud Infrastructure / Distributed Systems / GPU Computing
Deliverable
infrastructure
Required skills
Distributed systems architecture, Kubernetes, Python, Systems programming (Go/C++/Rust/Java), Performance optimization, Multi-tenancy design, Observability
Preferred skills
GPU scheduling, ML framework internals (PyTorch, FSDP, DeepSpeed), Inference stacks (vLLM, TensorRT-LLM, Triton, Ray Serve)
Technologies
Kubernetes, PyTorch, FSDP, DeepSpeed, vLLM, TensorRT-LLM, Triton, Ray Serve
Responsibilities
Own architecture and roadmap for ML compute/inference components; Design distributed systems for large-scale training and inference; Drive multi-tenancy and cost efficiency across GPU fleets; Build model deployment paths from experiment to production; Set engineering standards for reliability and performance; Partner with researchers to inform capacity strategy; Provide technical leadership and mentorship
Seniority
Staff, hands-on IC with strategic direction