CareerPlanGet AI match score →

Inference Engineer

*HQ - San Francisco, CA💼 Full-time🗓 2024-12-12 → 2026-07-31

Core

Design and build low latency, scalable, and reliable model inference and serving stack for cutting-edge foundation models including Transformers, SSMs, and hybrid models.

Role type

Senior Inference Engineer (Systems & ML)

Builds

Real-time multimodal intelligence products and inference infrastructure

Domain

Generative AI, Large Language Models, State Space Models

Deliverable

production ML models

Required skills

Distributed systems engineering, Inference pipeline design, Machine Learning model implementation, Generative AI experience, CUDA programming, Triton, vLLM, SGLang, Continuous Batching

Preferred skills

Experience with hybrid models, Zero-to-one execution

Technologies

Transformers, SSMs, vLLM, SGLang, CUDA, Triton

Responsibilities

Design and build low latency, scalable, and reliable model inference and serving stack; Work closely with research and product engineers to serve products in a fast, cost-effective, and reliable manner; Design and build robust inference infrastructure and monitoring for products.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗