CareerPlanSign in

GPU Optimization Engineer

San Francisco, CA💼 Full-time💰 $300,000–$300,000🗓 2026-02-11 → 2026-09-26

Core

Building low-latency AI systems for real-time speech and multimodal workloads, targeting sub-50ms time-to-first-token at 100+ concurrent requests on H100 GPUs.

Role type

Senior IC GPU Optimization Engineer (Inference)

Builds

Production inference runtimes for large generative models (speech/multimodal)

Domain

AI/ML Inference, GPU Architecture, Real-time Systems

Deliverable

production ML models

Required skills

CUDA programming, Triton kernel development, GPU memory hierarchy optimization, kernel fusion, attention mechanism optimization, KV cache management, low-level profiling tools, vLLM-style system modification

Preferred skills

Experience with AMD accelerators, model quantization techniques, decoding path optimization

Technologies

CUDA, Triton, vLLM, NVIDIA H100, AMD accelerators

Responsibilities

Profile GPU bottlenecks across memory bandwidth, kernel fusion, and scheduling; Write and tune custom CUDA/Triton kernels for performance-critical paths; Improve attention, decoding, and KV cache efficiency in inference runtimes; Modify and extend vLLM-style systems for real-time workloads; Optimize models to fit GPU memory constraints without degrading quality; Benchmark across NVIDIA and AMD GPUs; Partner with research to integrate new model ideas into production

Seniority

Senior, hands-on IC

Sourced via techire · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.