CareerPlanGet AI match score →

Sr. AI Inference Systems Engineer

US-California-Palo Alto💼 Full-time💰 $124,800–$124,800🗓 2026-07-02 → 2026-07-30

Core

Lead end-to-end optimization of the full inference pipeline for Large Models (LLM, Multimodal) to maximize throughput and minimize latency.

Role type

Senior IC AI Inference Systems Engineer

Builds

High-performance inference frameworks and optimized inference pipelines for cloud and smart industry solutions

Domain

Cloud computing, AI inference, heterogeneous computing

Deliverable

production ML models

Required skills

AI inference optimization, heterogeneous computing, KV Cache management, Quantization, Intelligent Routing, parallel computing, distributed systems, low-level programming (CUDA, Triton), deep learning frameworks (PyTorch, TensorFlow)

Preferred skills

Ultra-large-scale model optimization, inference cluster tuning, AI inference productization, technical publications or patents

Technologies

CUDA, Triton, PyTorch, TensorFlow

Responsibilities

Design and implement high-performance inference frameworks; optimize scheduling and memory management; conduct research on hardware accelerators; lead efforts to overcome technical bottlenecks; mentor team members

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗