Staff Software Engineer, Deep Learning Acceleration
Core
Optimizing Deep Learning network performance for Autonomous Vehicle systems, focusing on reducing latency and maximizing throughput in both onboard execution and large-scale data center training.
Role type
Staff Software Engineer (Deep Learning Acceleration)
Builds
High-performance inference and training pipelines for self-driving software
Domain
Autonomous Vehicles / Deep Learning / High-Performance Computing
Deliverable
production ML models
Required skills
CUDA, C++, Python, high-performance computing, parallel programming, GPU memory optimization, latency reduction, profiling (NVIDIA Nsight Systems/Compute), roofline model analysis, PyTorch or TensorFlow, computer vision fundamentals, transformer architectures
Preferred skills
TensorRT, OpenAI Triton, Mojo, motion planning, robotics, systems software
Technologies
NVIDIA Nsight Systems, NVIDIA Nsight Compute, PyTorch, TensorFlow, Linux/Unix
Responsibilities
Conduct performance analysis and optimization of Deep Learning networks on AVs; Optimize software architecture and latency for deep learning applications; Deploy deep learning models on AVs and train on large-scale data centers; Troubleshoot performance issues using profiling and roofline model techniques; Collaborate with cross-functional teams to enhance self-driving technology efficiency
Seniority
Staff, hands-on IC