Senior Accelerated Computing Architect
Core
Designing and optimizing software and system architectures for high-performance computing, scientific computing, machine learning, AI, datacenter, and automotive computing on NVIDIA GPUs.
Role type
Senior IC accelerated computing architect (hardware-software co-design) (via careerplan.io/jobs/893395048137-senior-accelerated-computing-architect-at-nvidia)
Builds
Optimized parallel algorithms, data structures, and reference codes for NVIDIA GPUs
Domain
Hardware-software co-design for accelerated computing (GPUs, AI, HPC)
Deliverable
production ML models
Required skills
GPU programming (CUDA/OpenCL), C/C++, parallel algorithms, performance optimization, hardware/software architecture analysis, benchmarking and profiling
Preferred skills
MPI/OpenSHMEM/NVSHMEM, threading APIs, Unix IPC, Python
Technologies
CUDA, OpenCL, MPI, OpenSHMEM, NVSHMEM, C, C++, Python
Responsibilities
Analyze and optimize performance on current/next-gen NVIDIA GPUs; Create and optimize core parallel algorithms and reference codes; Collaborate with hardware design, software engineering, product, and research teams; Facilitate software-hardware co-design in accelerated computing applications; Write white papers, conference publications, blog posts, and patent applications.
Seniority
Senior, hands-on IC
