CareerPlanSign in

Member of Technical Staff - ML Infra

San Francisco, CA💼 Full-time🗓 2026-09-20 → 2026-09-25

Core

Design, deploy, and maintain large distributed ML training and inference clusters; develop scalable pipelines for petabyte-scale datasets and model training.

Role type

Senior IC ML Infrastructure Engineer

Builds

Distributed ML training and inference clusters, scalable data pipelines, and model serving architectures

Domain

Cloud Infrastructure, Machine Learning Systems

Deliverable

production ML models

Required skills

Distributed training frameworks (FSDP, DeepSpeed), Cloud platforms (GCP, AWS, Azure), Containerization and orchestration (Kubernetes, Docker), GPU operations optimization, Scalable model serving architectures, Monitoring and observability

Preferred skills

Research on parallelization techniques, Numerical precision trade-offs, Distributed task management systems

Technologies

FSDP, DeepSpeed, Kubernetes, Docker, GCP, AWS, Azure

Responsibilities

Design and maintain distributed ML clusters, Develop end-to-end data and training pipelines, Research and test training approaches, Profile and debug low-level GPU operations, Integrate new research ideas into production

Seniority

Senior, hands-on IC

Sourced via dover · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.