Full Stack LLM Engineer
Core
Bringing up state-of-the-art open-source and proprietary ML models onto Cerebras CSX systems to achieve high performance, efficiency, and scalability.
Role type
Senior IC Full Stack LLM Engineer (Inference Bringup)
Builds
Production ML inference workloads on custom AI hardware
Domain
AI Hardware / Compiler Optimization / Deep Learning
Deliverable
production ML models
Required skills
C/C++ programming, compiler development (LLVM/MLIR), deep learning frameworks (PyTorch/TensorFlow), model internals (attention/MoE/diffusion), low-level optimization, performance profiling, debugging complex runtime issues
Preferred skills
Experience with NP-hard optimization problems, system-minded generalist approach
Technologies
Cerebras CSX, Python, LLVM, MLIR, PyTorch, TensorFlow
Responsibilities
Contribute to end-to-end bring up of ML models on Cerebras CSX systems; Work across the stack including model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning; Debug performance and correctness issues spanning model code, compiler IRs, runtime behavior, and hardware utilization; Propose and prototype improvements across tools, APIs, or automation flows to accelerate future bring ups
Seniority
Senior, hands-on IC