Senior Performance Engineer, Inference
Core
Design reproducible benchmarks for Cerebras AI chip inference performance and build a competitive pricing intelligence model to support sales and product strategy.
Role type
Senior Performance Engineer (Inference)
Builds
Reproducible inference benchmark suites and a living competitive pricing model for AI inference providers.
Domain
AI Hardware / Large Language Model Inference / High-Performance Computing
Deliverable
production ML models | dashboards & analysis | client delivery
Required skills
vLLM, SGLang, TensorRT-LLM, CUDA, Triton, Transformer architecture internals, KV-cache management, LLM inference economics, GPU memory hierarchies
Preferred skills
ML research background, open-source inference contributions, kernel optimization experience
Technologies
vLLM, SGLang, TensorRT-LLM, CUDA, Triton
Responsibilities
Design standardized benchmark suites for inference workloads; Evaluate new kernel fusions and quantization techniques; Build and update a competitive pricing model; Synthesize industry findings into actionable briefs for Sales and Product; Partner with Sales to build deal-specific competitive analyses; Collaborate with Engineering to identify competitive gaps.
Seniority
Senior, hands-on IC