CareerPlanSign in

ML Runtime and Kernel Engineer - Core ML

Sunnyvale, CA💼 Full-time🗓 2026-09-28 → 2026-09-29

Core

Develop novel machine learning algorithms and efficient execution on Cerebras Wafer-Scale Engine systems, bridging research ideas with high-performance runtime and kernel implementations.

Role type

Senior IC machine learning systems engineer (runtime & kernels)

Builds

High-performance ML runtimes, compilers, and low-level kernels for LLM training and inference on Cerebras hardware

Domain

AI hardware systems, machine learning infrastructure, high-performance computing

Deliverable

production ML models

Required skills

C++, Python, parallel programming, memory management, concurrency, performance optimization, system profiling and debugging, ML frameworks (PyTorch/JAX)

Preferred skills

CUDA, Triton, low-level assembly, compiler internals, distributed runtimes, HPC systems, LLM training/inference (attention, KV-cache), open-source contributions

Technologies

Cerebras Wafer-Scale Engine, PyTorch, JAX, CUDA, Triton

Responsibilities

Design and implement runtime components and high-performance kernels; Translate research prototypes into efficient implementations; Profile and debug performance across framework, compiler, runtime, and kernel layers; Optimize computation, memory movement, and concurrency; Develop benchmarks and automated tests; Collaborate with researchers and engineers on design alternatives; Contribute to software architecture and roadmap decisions

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.