Senior Machine Learning Engineer (Inference Platform)
Core
Own the end-to-end lifecycle of production ML serving systems for a live conversational shopping agent, focusing on inference infrastructure, reliability, and scaling.
Role type
Senior IC machine learning engineer (inference platform)
Builds
Multi-engine inference platform supporting LLMs, embedding models, and extraction models for an AI shopping agent
Domain
E-commerce / AI Shopping Agents / Large Language Model Serving
Deliverable
production ML models
Required skills
LLM serving engine expertise (vLLM, TGI, TensorRT-LLM, SGLang), Python, systems and infrastructure knowledge, cloud platforms (AWS, GCP, Azure), inference performance optimization (continuous batching, KV-cache, quantization), heterogeneous workload management, CI/CD integration
Preferred skills
High-growth startup experience, fast-moving technical landscape adaptability
Technologies
vLLM, TGI, TensorRT-LLM, SGLang, Python, AWS, GCP, Azure
Responsibilities
Own and evolve multi-engine inference platform, build production ML pipelines, define model versioning and lifecycle management strategies, enforce serving-layer SLAs (latency, availability, GPU utilization), build observability and monitoring tooling, optimize inference performance and resource utilization, partner with cross-functional teams on technical decisions
Seniority
Senior, hands-on IC