CareerPlanSign in

AI Inference Engineer

San Jose, US💼 Full-time💰 $176,600–$176,600🗓 2026-08-27 → 2026-09-25

Core

Optimizing Large Language Models (LLMs) for inference across diverse environments from data centers to edge devices, focusing on maximizing throughput and minimizing latency.

Role type

Senior IC AI Inference Engineer

Builds

High-performance inference engines and scalable AI serving infrastructure

Domain

AI/ML Infrastructure, High-Performance Computing, Cloud Infrastructure

Deliverable

production ML models

Required skills

Python, C++, Rust, Golang, vLLM, TensorRT, Llama.cpp, Ollama, Docker, Kubernetes, AWS, GCP, Azure, NVIDIA GPU optimization, TPU optimization

Preferred skills

Speculative Decoding, PagedAttention, open-source inference library contributions, CUDA kernel development, MLOps, SRE

Technologies

vLLM, TGI, NVIDIA Triton, TensorRT, Llama.cpp, Ollama, CUDA, CoreML, Kubernetes, Docker

Responsibilities

Build and maintain robust inference engines; Profile and optimize models for specialized hardware backends; Design and implement auto-scaling architectures for inference pipelines; Establish robust observability frameworks for performance monitoring.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.