Senior System Software Engineer
Core
Building efficient on-device AI software for high-performance local inference, low latency, and efficient memory use on resource-constrained platforms.
Role type
Senior IC systems software engineer (local AI inference)
Builds
Modern inference runtimes and execution stacks for LLMs, vision-language, TTS, ASR, and diffusion AI workloads on RTX and DGX GPUs.
Domain
AI computing, GPU systems, local inference
Deliverable
production ML models
Required skills
C++ programming, data structures, algorithms, machine learning, AI inferencing pipelines, CUDA, system-level debugging, performance optimization, quantization, pruning, sparsity, distillation, memory management, graph execution, hardware-aware optimization
Preferred skills
Generative AI, open-source contributions, distributed team delivery, Vulkan, DirectX, vLLM
Technologies
Llama.cpp, vLLM, PyTorch, WinML, DXCGC, TensorRT-RTX, CUDA, Vulkan, DirectX
Responsibilities
Partner with software, research, architecture, and product teams to align strategies and technical needs; Build and optimize local AI inference stack for RTX, RTX Pro, and DGX GPUs; Architecture and development of modern inference runtimes and execution stacks; Perform end-to-end optimization of AI models, data pipelines, and inference runtimes; Perform system-level debugging, performance optimization, and performance–accuracy trade-off analysis.
Seniority
Senior, hands-on IC