ML Research Engineer (Inference)
Core
Adapt advanced language and vision models to run efficiently on Cerebras' flagship AI architecture, focusing on speculative decoding, pruning, compression, and sparse attention to deliver low-latency, high-throughput inference.
Role type
Research Engineer (Inference ML)
Builds
Optimized inference workloads for Cerebras hardware
Domain
AI hardware acceleration / Generative AI
Deliverable
production ML models
Required skills
Python, C++, PyTorch, Transformers, vLLM, SGLang, deep learning concepts, neural networks, transformers, Generative AI systems
Preferred skills
speculative decoding, neural network pruning and compression, sparse attention, quantization, sparsity, post-training techniques, inference-focused evaluations, Linux environments
Technologies
Cerebras architecture, PyTorch, Hugging Face Transformers, vLLM, SGLang
Responsibilities
Implement and adapt transformer-based models (NLP and/or vision) to run on Cerebras hardware, optimize models for inference performance (latency, throughput), run experiments and analyze results, bring up and validate models on the Cerebras system, debug and troubleshoot model or system issues, support profiling and performance analysis using internal tools, collaborate with cross-functional teams (ML, software, hardware) on model integration
Seniority
Mid-level, hands-on IC