Staff Software Engineer, Deep Learning Acceleration
Core
Optimizing deep learning network performance and latency for autonomous vehicle systems and large-scale data center training.
Role type
Staff Software Engineer (Deep Learning Acceleration)
Builds
High-performance inference and training pipelines for self-driving technology
Domain
Autonomous Vehicles / Deep Learning / High-Performance Computing
Deliverable
production ML models
Required skills
CUDA, C++, Python, high-performance computing, parallel programming, GPU memory optimization, latency reduction, profiling (NVIDIA Nsight Systems/Compute), roofline model analysis, deep learning frameworks (PyTorch, TensorFlow), computer vision fundamentals, transformer architectures
Preferred skills
motion planning, robotics, systems software, TensorRT, OpenAI Triton, Mojo
Technologies
NVIDIA Nsight Systems, NVIDIA Nsight Compute, PyTorch, TensorFlow, TensorRT, OpenAI Triton, Mojo
Responsibilities
Conduct performance analysis and optimization of Deep Learning networks on AVs; Optimize software architecture and latency for deep learning applications; Deploy deep learning models on AVs and train on large-scale data centers; Troubleshoot performance issues using profiling and roofline model techniques; Collaborate with cross-functional teams to enhance self-driving technology efficiency
Seniority
Staff, hands-on IC