Director/Sr. Manager, AI Inference Model Scaling
Core
Define technical vision and strategy for enabling state-of-the-art foundation models and generative AI workloads on Cerebras' Wafer-Scale Engine (WSE) hardware.
Role type
Director/Senior Manager, AI Inference Model Scaling
Builds
Compiler frontend, model transformation pipeline, graph optimization infrastructure, high-performance kernels, and runtime integration for next-generation AI models.
Domain
AI hardware architecture, machine learning frameworks, compiler technologies, distributed systems, and model optimization.
Deliverable
production ML models
Required skills
Compiler infrastructure design, graph compilation and optimization, Python, C++, production software delivery, organizational strategy, team leadership, cross-functional collaboration.
Preferred skills
Compiler frontends for AI accelerators, PyTorch/JAX/TensorFlow/ONNX support, LLM inference systems, distributed compilation, hardware/software co-design.
Technologies
LLVM, MLIR, XLA, TVM, Torch FX, PyTorch, JAX, TensorFlow, ONNX, WSE.
Responsibilities
Define technical roadmap and strategy; establish technical direction across multiple teams; lead design reviews and engineering standards; hire, mentor, and grow engineering leaders; drive organizational planning and headcount strategy; partner with Cloud Platform, ML, and Hardware teams; collaborate with Product Management and customers on model bring-up; balance rapid model support with long-term compiler architecture.
Seniority
Senior, hands-on IC with organizational leadership