Member of Technical Staff, Performance Optimization
Core
Software Engineer optimizing performance and efficiency across AI infrastructure layers, from GPU kernels to distributed systems, for large-scale generative AI models.
Role type
Senior IC performance optimization engineer (AI infrastructure)
Builds
High-throughput AI workloads including LLMs, VLMs, and video models
Domain
Artificial Intelligence / High-Performance Computing / GPU Systems
Deliverable
production ML models
Required skills
CUDA, ROCm, GPU profiling (Nsight, nvprof, CUPTI), PyTorch, distributed system debugging, GPU architecture, parallel programming, compute kernels
Preferred skills
ML compilers (torch.compile, Triton, XLA), open-source ML/HPC contributions, cloud-scale orchestration (Kubernetes), hardware-aware model design
Technologies
CUDA, Triton, PyTorch, Nsight, nvprof, CUPTI, Kubernetes, RDMA, Infiniband, RoCE
Responsibilities
Optimize system and GPU performance for high-throughput AI workloads; Analyze and improve latency, throughput, memory usage, and compute efficiency; Profile system performance to detect and resolve GPU- and kernel-level bottlenecks; Implement low-level optimizations using CUDA, Triton, and other performance tooling; Drive improvements in execution speed and resource utilization for large-scale model workloads; Collaborate with ML researchers to co-design and tune model architectures for hardware efficiency; Build and maintain performance benchmarking and monitoring infrastructure; Scale inference and training systems across multi-GPU, multi-node environments; Evaluate and integrate optimizations for emerging hardware accelerators and specialized runtimes
Seniority
Senior, hands-on IC