Senior Deep Learning Frameworks CUDA Software Engineer
Core
Integrate new CUDA features and Runtime abstractions into AI frameworks (PyTorch, TRT-LLM, vLLM, SGLang, JAX) to enable scaling Deep Learning and HPC applications from microsecond latency inference to 100K GPU training clusters.
Role type
Senior IC CUDA software engineer (Deep Learning frameworks)
Builds
AI toolkits, production ML models, distributed runtime systems, and open-source framework integrations
Domain
High Performance Computing (HPC) and Artificial Intelligence
Deliverable
production ML models
Required skills
CUDA development, Deep Learning Frameworks (PyTorch, JAX), C++, Python, AI Compiler-Runtime interface design, fault-tolerant system design, performance benchmarking, HPC/AI communication concepts, computer system architecture
Preferred skills
Deep learning compilers (Triton, XLA, torch.compile), distributed machine learning techniques (pipeline/tensor parallelism), kernel authoring (cuTe), compute & communication overlap programming
Technologies
CUDA, PyTorch, JAX, TRT-LLM, vLLM, SGLang, NCCL, MPI, UCX, Triton, Nsight Systems
Responsibilities
Integrate CUDA features from PoC to production, analyze AI workloads to identify lower-layer innovation opportunities, drive improvements in AI Compiler-Runtime interface, design elastic solutions for large-scale workloads, influence core CUDA roadmap, develop exploratory profiling tools, write maintainable code for open-source and commercial products
Seniority
Senior, hands-on IC