Compute Architecture Software Engineer
Core
Develop and optimize software solutions to accelerate LLM inference using GPU technology for NVIDIA's TRTLLM project.
Role type
Senior IC LLM inference software engineer (GPU)
Builds
Optimized LLM inference software running on single PCs to multi-GPU clusters
Domain
AI / Deep Learning / GPU Computing
Deliverable
production ML models
Required skills
GPU programming, LLM inference, Python, C++, CUDA, deep learning frameworks
Preferred skills
None stated
Technologies
CUDA, Python, C++, deep learning frameworks
Responsibilities
Develop and optimize software to accelerate LLM inference using GPU technology; Collaborate with engineers to implement and refine GPU-based algorithms; Analyze methods to improve performance across diverse computing environments
Seniority
Senior, hands-on IC
Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
