Senior AI/ML Performance Engineer
Core
Designing and optimizing large-scale GPU infrastructure and ML training/inference environments for autonomous vehicle development.
Role type
Senior IC performance engineer (ML infrastructure)
Builds
Scalable GPU clusters and high-performance computing environments for AV model training and inference
Domain
Automotive (AV) + High-Performance Computing (HPC)
Deliverable
infrastructure
Required skills
Python, PyTorch, Kubernetes, distributed systems, HPC, GPU monitoring (Nvidia DCGM, nvidia-smi, Grafana), cloud platforms (AWS/GCP/Azure)
Preferred skills
Enterprise Nvidia GPU architectures (H100, B200, GB200), Hugging Face model deployment, BigQuery, Nvidia Nsight profiling
Technologies
PyTorch, Kubernetes, Nvidia DCGM, nvidia-smi, Grafana, AWS, GCP, Azure, Hugging Face, BigQuery, Nvidia Nsight
Responsibilities
Adopt and run AV models to support long-term GPU system strategy; conduct deep-dive analyses of production workloads to identify bottlenecks; partner with AI/ML Research and Cloud Vendors to enhance engineering velocity; identify opportunities for architectural improvements to ensure scalability and reliability of large-scale ML environments
Seniority
Senior, hands-on IC