GPU Optimization Engineer
Core
Building low-latency AI systems for real-time speech and multimodal workloads, targeting sub-50ms time-to-first-token at 100+ concurrent requests on H100 GPUs.
Role type
Senior IC GPU Optimization Engineer (Inference)
Builds
Production inference runtimes for large generative models (speech/multimodal)
Domain
AI/ML Inference, GPU Architecture, Real-time Systems
Deliverable
production ML models
Required skills
CUDA programming, Triton kernel development, GPU memory hierarchy optimization, kernel fusion, attention mechanism optimization, KV cache management, low-level profiling tools, vLLM-style system modification
Preferred skills
Experience with AMD accelerators, model quantization techniques, decoding path optimization
Technologies
CUDA, Triton, vLLM, NVIDIA H100, AMD accelerators
Responsibilities
Profile GPU bottlenecks across memory bandwidth, kernel fusion, and scheduling; Write and tune custom CUDA/Triton kernels for performance-critical paths; Improve attention, decoding, and KV cache efficiency in inference runtimes; Modify and extend vLLM-style systems for real-time workloads; Optimize models to fit GPU memory constraints without degrading quality; Benchmark across NVIDIA and AMD GPUs; Partner with research to integrate new model ideas into production
Seniority
Senior, hands-on IC