Senior Software Engineer - GPU Kernel Authoring & Optimization
Core
Senior GPU kernel authoring and optimization engineer for large-scale LLM inference stacks.
Role type
Senior IC GPU systems engineer
Builds
High-performance GPU kernels and benchmarking workflows for model-serving stacks
Domain
AI/ML infrastructure, GPU computing, LLM inference
Deliverable
production ML models
Required skills
CUDA kernel authoring, GPU architecture (tensor cores, memory hierarchy), C++, Python, performance profiling (Nsight), model-serving stacks (vLLM, TensorRT-LLM)
Preferred skills
Triton, Mojo, CuTe DSL, JAX/Pallas, HIP/ROCm, NCCL, Kubernetes, Slurm, MLPerf, OSS contributions
Technologies
CUDA, C++, Python, Nsight Compute, Nsight Systems, vLLM, TensorRT-LLM, llm-d, SGLang, Triton, Mojo, CuTe, JAX, Pallas, HIP, ROCm, NCCL, Kubernetes, Slurm
Responsibilities
Author and optimize CUDA kernels for GEMMs, attention, MoE routing, and quantization; tune occupancy and memory coalescing; build reproducible microbenchmarks and roofline analyses; implement MLPerf Inference/Training workflows; lead design reviews and mentor junior engineers
Seniority
Senior, hands-on IC