CareerPlanGet AI match score →

Machine Learning Engineer - Inference

San Francisco💼 Full-time💰 $160,000–$160,000🗓 2026-07-10 → 2026-07-31

Core

Design and build production systems to optimize and scale AI inference for large language models.

Role type

Senior IC machine learning engineer (inference systems)

Builds

High-performance inference engine services and runtime systems for large-scale AI applications

Domain

Artificial Intelligence / Large Language Models / Systems Engineering

Deliverable

production ML models

Required skills

Python, PyTorch, high-performance system design, multi-threading, memory management, networking, storage, code reviews, fault-tolerant system implementation

Preferred skills

TGI, vLLM, TensorRT-LLM, Optimum, speculative decoding, CUDA, Triton, Rust, Cython

Technologies

PyTorch, CUDA, Triton, Rust, Cython

Responsibilities

Design and build production systems for the inference engine; Develop and optimize runtime inference services; Collaborate with researchers and engineers to bring new features; Conduct design and code reviews; Create services, tools, and documentation; Implement robust and fault-tolerant systems for data ingestion and processing

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗