CareerPlanGet AI match score →

Senior Performance Engineer, Inference

Headquarters/Sunnyvale Office💼 Full-time🗓 2026-04-13 → 2026-07-31

Core

Design reproducible benchmarks for Cerebras AI chip inference performance and build a competitive pricing intelligence model to support sales and product strategy.

Role type

Senior Performance Engineer (Inference)

Builds

Reproducible inference benchmark suites and a living competitive pricing model for AI inference providers.

Domain

AI Hardware / Large Language Model Inference / High-Performance Computing

Deliverable

production ML models | dashboards & analysis | client delivery

Required skills

vLLM, SGLang, TensorRT-LLM, CUDA, Triton, Transformer architecture internals, KV-cache management, LLM inference economics, GPU memory hierarchies

Preferred skills

ML research background, open-source inference contributions, kernel optimization experience

Technologies

vLLM, SGLang, TensorRT-LLM, CUDA, Triton

Responsibilities

Design standardized benchmark suites for inference workloads; Evaluate new kernel fusions and quantization techniques; Build and update a competitive pricing model; Synthesize industry findings into actionable briefs for Sales and Product; Partner with Sales to build deal-specific competitive analyses; Collaborate with Engineering to identify competitive gaps.

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Ashby ↗