Member of Technical Staff, Performance Optimization
Core
Optimizing performance, latency, and efficiency for large-scale generative AI workloads (LLMs, VLMs, video models) across GPU kernels and distributed systems.
Role type
Senior IC performance optimization engineer (GPU/HPC)
Builds
High-throughput inference and training systems for generative AI models
Domain
Generative AI infrastructure, High-Performance Computing (HPC), GPU systems
Deliverable
production ML models
Required skills
CUDA/ROCm, GPU profiling (Nsight, nvprof, CUPTI), PyTorch, distributed system debugging, GPU architecture, parallel programming, compute kernels
Preferred skills
ML compilers (torch.compile, Triton, XLA), open-source ML/HPC contributions, Kubernetes, hardware-aware model design
Technologies
CUDA, Triton, PyTorch, Nsight, nvprof, CUPTI, Infiniband, RoCE
Responsibilities
Optimize system and GPU performance for high-throughput AI workloads; Analyze and improve latency, throughput, memory usage, and compute efficiency; Profile system performance to detect and resolve GPU- and kernel-level bottlenecks; Implement low-level optimizations using CUDA, Triton, and other performance tooling; Drive improvements in execution speed and resource utilization for large-scale model workloads; Collaborate with ML researchers to co-design and tune model architectures for hardware efficiency
Seniority
Senior, hands-on IC