CareerPlanSign in

AI Systems Research and Development Engineer – LLM Inference Systems & Optimization

US-WA-Bellevue💼 Full-time🗓 2026-09-09 → 2026-09-25

Core

Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.

Role type

Senior IC systems engineer (LLM inference optimization)

Builds

Next-generation high-performance and intelligent inference systems for agentic enterprise workloads

Domain

AI Systems / Large Language Model Inference / High-Performance Computing

Deliverable

production ML models

Required skills

LLM inference system design, distributed AI systems, GPU systems, high-performance computing, CUDA/Triton programming, performance profiling (Nsight), system optimization, parallel decoding strategies, KV-cache management, model-system co-design

Preferred skills

AI-native engineering approaches, automated profiling and configuration search, open-source contribution

Technologies

vLLM, SGLang, TensorRT-LLM, CUTLASS, cuBLAS, cuDNN, Nsight Systems, Nsight Compute

Responsibilities

Design distributed inference strategies across GPUs and nodes; Develop efficient approaches for multi-model serving and dynamic resource management; Analyze and optimize GPU kernels and operators; Profile and benchmark end-to-end workloads to identify bottlenecks; Collaborate with model researchers and infrastructure teams to deploy innovations; Open-source and publish innovations.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.