Research Engineer, Infrastructure, Inference
Core
Design, optimize, and scale infrastructure systems to enable high-performance, cost-effective, and reliable inference for large AI models.
Role type
Senior IC infrastructure research engineer (AI inference systems)
Builds
Scalable inference serving systems, orchestration frameworks, and compute fleets for AI models
Domain
Artificial Intelligence / Large Language Model Infrastructure
Deliverable
production ML models
Required skills
Deep learning frameworks (PyTorch, JAX), inference serving systems (SGLang, vLLM), distributed compute systems, GPU parallelism, hardware-aware optimizations, orchestration frameworks (Kubernetes, Ray, SLURM), codebase optimization, observability standards
Preferred skills
Experience with large-scale language models (hundreds of billions of parameters), open-source ML/systems contributions, improving research productivity via infrastructure design
Technologies
PyTorch, JAX, SGLang, vLLM, Kubernetes, Ray, SLURM, Triton, DeepSpeed, XLA
Responsibilities
Design and implement new techniques/tools/architectures to improve performance, latency, throughput, and efficiency; Optimize codebase and compute fleet (GPUs) to utilize hardware FLOPs, bandwidth, and memory; Extend orchestration frameworks for distributed inference, evaluation, and large-batch serving; Establish standards for reliability, observability, and reproducibility across the inference stack
Seniority
Senior, hands-on IC