Research Scientist / Engineer – Performance Optimization
Core
Profile and optimize GPU/CPU/accelerator code to make multimodal models train efficiently and deploy at scale.
Role type
Senior IC machine-learning engineer (performance optimization)
Builds
High-performance training and inference pipelines for multimodal models
Domain
AI/ML, GPU computing, distributed systems
Deliverable
production ML models
Required skills
CUDA programming, Triton kernel development, PyTorch kernel development, GPU profiling, transformer architecture internals, distributed multi-node deployment
Preferred skills
Compiler optimization (torch.compile, TensorRT, ONNX, XLA), inference latency optimization, warp-level intrinsics
Technologies
CUDA, Triton, PyTorch, NVIDIA Nsight, torch profiler
Responsibilities
Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency; Write high-performance PyTorch, Triton, and CUDA kernels; Develop fused kernels leveraging tensor cores; Optimize model architectures for distributed multi-node production deployment; Build performance monitoring and analysis tools; Research and implement cutting-edge optimization techniques for transformer models
Seniority
Senior, hands-on IC
