Staff Kernel Optimzation Engineer
Core
Develop high-performance software solutions at the intersection of hardware and software, implementing and optimizing deep learning operations for custom massively parallel processor architecture.
Role type
Staff Kernel Optimization Engineer
Builds
High-performance ML and HPC kernels, parallel and distributed algorithms for Cerebras WSE System
Domain
AI hardware acceleration, High-Performance Computing (HPC), Custom Silicon
Deliverable
production ML models
Required skills
C++, Python, hardware architecture concepts, low-level assembly programming, custom C-like domain specific language (CSL), debugging complex software stacks, mathematical modeling and analysis, unit and system testing methodologies
Preferred skills
kernel development and testing, parallel algorithms, distributed memory systems, accelerator programming (GPUs, FPGAs), Machine Learning neural networks frameworks (TensorFlow, PyTorch), HPC kernels optimization
Technologies
Cerebras WSE System, TensorFlow, PyTorch, GPUs, FPGAs
Responsibilities
Develop design specifications for new machine learning and linear algebra kernels; Develop and debug kernel library of highly optimized low level assembly instruction and C-like domain specific language routines; Develop and debug high-performance kernel routines in low-level assembly and a custom C-like (CSL) language; Using mathematical models and analysis to measure the software performance and inform design decisions; Develop and integrate unit and system testing methodologies to verify correct functionality and performance of kernel libraries; Study emerging trends in Machine Learning applications and help evolve Kernel library architecture to address computational challenges of the start-of-the-art Neural Networks; Interact with chip and system architects to optimize instruction sets, microarchitecture, and IO of next generation systems.
Seniority
Staff, hands-on IC with strategic impact