CareerPlanGet AI match score →

Member of Technical Staff, Performance Optimization

San Mateo💼 Full-time💰 $175,000–$175,000🗓 2026-06-05 → 2026-07-31

Core

Optimizing performance, latency, and efficiency for large-scale generative AI workloads (LLMs, VLMs, video models) across GPU kernels and distributed systems.

Role type

Senior IC performance optimization engineer (GPU/HPC)

Builds

High-throughput inference and training systems for generative AI models

Domain

Generative AI infrastructure, High-Performance Computing (HPC), GPU systems

Deliverable

production ML models

Required skills

CUDA/ROCm, GPU profiling (Nsight, nvprof, CUPTI), PyTorch, distributed system debugging, GPU architecture, parallel programming, compute kernels

Preferred skills

ML compilers (torch.compile, Triton, XLA), open-source ML/HPC contributions, Kubernetes, hardware-aware model design

Technologies

CUDA, Triton, PyTorch, Nsight, nvprof, CUPTI, Infiniband, RoCE

Responsibilities

Optimize system and GPU performance for high-throughput AI workloads; Analyze and improve latency, throughput, memory usage, and compute efficiency; Profile system performance to detect and resolve GPU- and kernel-level bottlenecks; Implement low-level optimizations using CUDA, Triton, and other performance tooling; Drive improvements in execution speed and resource utilization for large-scale model workloads; Collaborate with ML researchers to co-design and tune model architectures for hardware efficiency

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗