Senior Deep Learning Software Infrastructure Engineer
Core
Build and scale training libraries and infrastructure enabling end-to-end autonomous driving models on thousands of GPUs.
Role type
Senior Deep Learning Software Infrastructure Engineer
Builds
Training libraries, orchestration frameworks, and fault-resilient systems for massive GPU clusters
Domain
Autonomous Vehicles / High-Performance Computing / Deep Learning Infrastructure
Deliverable
production ML models
Required skills
Distributed systems design, Deep learning frameworks (PyTorch), Large-scale training (DDP/FSDP, NCCL, tensor/pipeline parallelism), Datacenter networking (RoCE, IB), Parallel filesystems (Lustre), Python production library development
Preferred skills
Scaling clusters with >1,000 GPUs, Fault resilience and high availability, Elastic training, Large-scale observability
Technologies
PyTorch, DDP, FSDP, NCCL, RoCE, IB, Lustre, Slurm, Kubernetes
Responsibilities
Craft and harden deep learning infrastructure libraries for multi-thousand GPU clusters; Improve efficiency in data loaders, distributed training, and scheduling; Build robust pipelines for massive video datasets; Collaborate with researchers to minimize training stalls; Own core infrastructure components like orchestration and fault-resilient systems; Partner with leadership to scale infrastructure with growing GPU capacity
Seniority
Senior, hands-on IC