CareerPlanSign in

Senior Deep Learning Software Infrastructure Engineer

US, CA, Remote💼 Full-time💰 $224,000–$224,000🗓 2026-09-07 → 2026-09-25

Core

Build and scale training libraries and infrastructure enabling end-to-end autonomous driving models on thousands of GPUs.

Role type

Senior Deep Learning Software Infrastructure Engineer

Builds

Training libraries, orchestration frameworks, and fault-resilient systems for massive GPU clusters

Domain

Autonomous Vehicles / High-Performance Computing / Deep Learning Infrastructure

Deliverable

production ML models

Required skills

Distributed systems design, Deep learning frameworks (PyTorch), Large-scale training (DDP/FSDP, NCCL, tensor/pipeline parallelism), Datacenter networking (RoCE, IB), Parallel filesystems (Lustre), Python production library development

Preferred skills

Scaling clusters with >1,000 GPUs, Fault resilience and high availability, Elastic training, Large-scale observability

Technologies

PyTorch, DDP, FSDP, NCCL, RoCE, IB, Lustre, Slurm, Kubernetes

Responsibilities

Craft and harden deep learning infrastructure libraries for multi-thousand GPU clusters; Improve efficiency in data loaders, distributed training, and scheduling; Build robust pipelines for massive video datasets; Collaborate with researchers to minimize training stalls; Own core infrastructure components like orchestration and fault-resilient systems; Partner with leadership to scale infrastructure with growing GPU capacity

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.