CareerPlanGet AI match score →

Staff Inference ML Runtime Engineer

US and Canada Offices💼 Full-time🗓 2025-11-25 → 2026-07-31

Core

Design and implement APIs, machine learning features, and tools enabling state-of-the-art generative AI models to run efficiently on custom hardware clusters.

Role type

Senior IC inference ML runtime engineer

Builds

Scalable, high-performance inference solutions for generative AI models

Domain

AI hardware acceleration / Large Language Model inference

Deliverable

production ML models

Required skills

Python, C++, multi-threaded programming, performance optimization, software architectural patterns, LLM serving frameworks, PyTorch, observability, automated testing

Preferred skills

Experience with vLLM, SGLang, TensorRT-LLM, multimodal model inference

Technologies

Python, C++, PyTorch, vLLM, SGLang, TensorRT-LLM

Responsibilities

Design and implement ML features for generative AI inference, maintain scalable serving backend, optimize software for high throughput and low latency, analyze and improve service efficiency, build robust automated test suites, lead cross-functional initiatives

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗