Senior ML Engineer (Token Factory)
Core
Building a high-performance inference and fine-tuning platform to maximize throughput, minimize latency, and optimize cost-per-token across tens of thousands of GPUs.
Role type
Senior IC machine-learning engineer (inference optimization & low-precision training)
Builds
High-performance inference engines and low-precision training pipelines for foundation models
Domain
Cloud infrastructure, Large Language Models (LLMs), GPU compute
Deliverable
production ML models
Required skills
Proficiency in Python, deep understanding of transformer architecture, experience profiling GPU workloads, knowledge of GPU memory hierarchy, familiarity with LLM concepts (MHA, RoPE, KV-cache, Flash Attention, quantization), understanding of large neural network training performance, strong software engineering skills, CI/CD proficiency
Preferred skills
Experience with open-source inference engines (vLLM, SGLang, TensorRT-LLM), proficiency in kernel languages (Triton, CUTLASS, CUDA), track record of building distributed systems or high-load web services
Technologies
Python, PyTorch, Nsight, vLLM, SGLang, TensorRT-LLM, Triton, CUDA
Responsibilities
Identify LLM inference bottlenecks to drive production speedups, implement novel speculative decoding architectures, optimize components of various LLM designs (dense/MoE, autoregressive/parallel), design and productionize low-precision (FP8, NVFP4/MXFP4) training and inference pipelines
Seniority
Senior, hands-on IC