Applied AI/ML Engineer
Core
Ship AI-powered products end-to-end on a GPU inference fleet, owning serving, reinforcement-learning post-training pipelines, and evaluation.
Role type
Senior Applied AI/ML Engineer (LLM inference & RL)
Builds
Production LLM inference services and RL/post-training pipelines
Domain
AI Infrastructure / GPU Compute / Large Language Models
Deliverable
production ML models
Required skills
LLM inference serving, reinforcement learning, post-training methods, Python, PyTorch, GPU execution optimization
Preferred skills
Distributed training, quantization, speculative decoding, Kubernetes, GPU fleet orchestration
Technologies
vLLM, SGLang, TensorRT-LLM, Megatron-LM, Prime Intellect, Ray, SkyPilot, Slurm
Responsibilities
Deploy and optimize LLM inference across the fleet; Build and operate RL and post-training pipelines; Build evaluation harnesses to measure quality, throughput, and cost; Partner with Infrastructure and Product teams
Seniority
Senior, hands-on IC