Research Engineer
Core
Optimize large-scale training and inference workloads on AI infrastructure to improve performance, scalability, and reliability for real-world AI systems.
Role type
Senior IC Research Engineer (ML Systems & Infrastructure)
Builds
Production AI workloads, inference pipelines, model serving systems, and performance-oriented tooling
Domain
AI Infrastructure / ML Systems / High-Performance Computing
Deliverable
production ML models
Required skills
Deep learning frameworks (PyTorch), large-scale training/inference workloads, distributed systems and parallelism strategies, software engineering fundamentals (API design, tooling, debugging), performance bottleneck analysis
Preferred skills
Inference optimization techniques (quantization, speculative decoding, mixed precision), CUDA/Triton/TensorRT/vLLM/SGLang/Dynamo, open-source contributions, startup experience, advanced degree in AI/ML/systems
Technologies
PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, Dynamo, NVIDIA GPUs, TPUs
Responsibilities
Optimize large-scale training and inference workloads across GPUs and distributed systems; Work with customers to analyze workloads and improve deployed AI system performance; Develop and improve inference pipelines and model serving systems; Design profiling, debugging, and observability tools; Partner with hardware vendors to support diverse compute backends; Contribute to open-source projects
Seniority
Senior, hands-on IC