Inference Performance & Deployment - Member of Technical Staff
Core
Design tooling and methodologies to ground heterogeneous AI infrastructure in real-world performance, acting as the integration point between engineering functions and production environments.
Role type
Member of Technical Staff (Inference Performance & Deployment)
Builds
Deployment patterns, benchmarking harnesses, regression suites, and performance dashboards for heterogeneous compute systems.
Domain
AI Infrastructure / Heterogeneous Compute / Inference Systems
Deliverable
production ML models
Required skills
Large model inference deployment, multi-node GPU deployment, end-to-end performance characterization, serving frameworks (Dynamo, Triton), system bottleneck isolation, reproducible measurement methodology
Preferred skills
Cloud instance self-hosting, orchestration and routing software, caching and request scheduling, resource allocation
Technologies
Cerebras, d-Matrix, Intel, NVIDIA, AMD, Normal Computing, Tenstorrent, GreatSky, Mixx, Dynamo, Triton Inference Server
Responsibilities
Run experiments self-hosting models across providers and hardware configurations, develop optimized deployment patterns, work on orchestration and routing software, act as integration point for accelerator support and engine features, build and maintain benchmarking harnesses and dashboards
Seniority
Mid-Senior, hands-on IC