Software Engineer, LLM Inference
Core
Develop robust, scalable inferencing software for LLMs and deep learning models, optimizing performance across multiple platforms.
Role type
Senior IC CPU computing engineer (LLM inference)
Builds
GPU-accelerated libraries (CUDA, cuDNN, TensorRT) and TensorRT Edge LLM features
Domain
AI-City and self-driving car solutions using GPU-accelerated computing
Deliverable
production ML models
Required skills
C/C++ programming, software design, performance analysis, debugging, test design, deep learning frameworks (PyTorch), LLM/generative model awareness
Preferred skills
Academic awareness of AI developments, proactive problem solving
Technologies
CUDA, cuDNN, TensorRT, TensorRT Edge, PyTorch
Responsibilities
Craft and develop robust inferencing software scaled to multiple platforms, perform performance analysis and optimization, update TensorRT and TensorRT Edge LLM based on academic developments, collaborate with software, research, and product teams
Seniority
Senior, hands-on IC
