Senior System Software Engineer, AI Infrastructure
Core
Evaluating performance and usability of NVIDIA's GPU-accelerated AI platforms, SDKs, and libraries through multi-node training/inference jobs, benchmarking, and optimization guidance.
Role type
Senior System Software Engineer (AI Infrastructure)
Builds
GPU-accelerated AI platforms, SDKs, libraries, and tools for generative AI, autonomous driving, and industrial robots.
Domain
AI Infrastructure, High-Performance Computing (HPC), GPU Computing
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, C++, multi-node cluster management (Slurm, Kubernetes), deep learning architectures, PyTorch, distributed training, CUDA, cuDNN, TensorRT-LLM, Triton, NCCL
Preferred skills
HPC cluster tuning, InfiniBand, NVLink, RoCE, RDMA, collective-comm libraries, modern LLM architectures, custom GPU kernels
Technologies
Slurm, Kubernetes, PyTorch, CUDA, cuDNN, TensorRT-LLM, Triton, NCCL, InfiniBand, NVLink, RoCE, RDMA
Responsibilities
Run multi-node training/inference jobs to assess performance and validate usability; Design benchmark suites for hardware, networking, and software stacks; Profile deep-learning workloads to identify bottlenecks and deliver optimization guidance; Produce tutorials, scripts, and whitepapers for customers and tech press; Analyze competitive solutions for product positioning; Present live demos at global conferences.
Seniority
Senior, hands-on IC