CareerPlanSign in

Research Scientist / Engineer – Performance Optimization

Redwood City, CA💼 Full-time🗓 2026-07-24 → 2026-09-26

Core

Profile and optimize GPU/CPU/accelerator code to make multimodal models train efficiently and deploy at scale.

Role type

Senior IC machine-learning engineer (performance optimization)

Builds

High-performance training and inference pipelines for multimodal models

Domain

AI/ML, GPU computing, distributed systems

Deliverable

production ML models

Required skills

CUDA programming, Triton kernel development, PyTorch kernel development, GPU profiling, transformer architecture internals, distributed multi-node deployment

Preferred skills

Compiler optimization (torch.compile, TensorRT, ONNX, XLA), inference latency optimization, warp-level intrinsics

Technologies

CUDA, Triton, PyTorch, NVIDIA Nsight, torch profiler

Responsibilities

Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency; Write high-performance PyTorch, Triton, and CUDA kernels; Develop fused kernels leveraging tensor cores; Optimize model architectures for distributed multi-node production deployment; Build performance monitoring and analysis tools; Research and implement cutting-edge optimization techniques for transformer models

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.