CareerPlanSign in

Senior Machine Learning Engineer, AI Platform

San Jose💼 Full-time💰 $172,500–$172,500🗓 2026-09-03 → 2026-09-26

Core

Design and operate distributed systems for large-scale GPU training and low-latency inference, managing multi-tenant fleets and moving models from experiment to production.

Role type

Staff Machine Learning Platform Engineer

Builds

Shared compute and inference platform for Adobe's AI products (Firefly, Creative Cloud, Experience Cloud)

Domain

Cloud Infrastructure / Distributed Systems / GPU Computing

Deliverable

infrastructure

Required skills

Distributed systems architecture, Kubernetes, Python, Systems programming (Go/C++/Rust/Java), Performance optimization, Multi-tenancy design, Observability

Preferred skills

GPU scheduling, ML framework internals (PyTorch, FSDP, DeepSpeed), Inference stacks (vLLM, TensorRT-LLM, Triton, Ray Serve)

Technologies

Kubernetes, PyTorch, FSDP, DeepSpeed, vLLM, TensorRT-LLM, Triton, Ray Serve

Responsibilities

Own architecture and roadmap for ML compute/inference components; Design distributed systems for large-scale training and inference; Drive multi-tenancy and cost efficiency across GPU fleets; Build model deployment paths from experiment to production; Set engineering standards for reliability and performance; Partner with researchers to inform capacity strategy; Provide technical leadership and mentorship

Seniority

Staff, hands-on IC with strategic direction

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.