Engineering Manager, Forward Deployed Engineering (LLM)
Core
Lead and mentor Forward Deployed Engineers to build, scale, and optimize LLM inference workloads for customers, acting as both a manager and a hands-on technical contributor.
Role type
Engineering Manager (Player & Coach)
Builds
Production-ready LLM inference systems, model servers, and custom AI workflows for enterprise customers.
Domain
Artificial Intelligence / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
Python, LLMs, inference optimization, serving frameworks (vLLM, TensorRT, Triton, Hugging Face, Ray Serve), observability, profiling, cost/performance tradeoffs
Preferred skills
GPU infrastructure, distributed inference, model compression, leading customer-facing engineering teams
Technologies
Python, vLLM, TensorRT, Triton, Hugging Face, Ray Serve, Docker
Responsibilities
Lead and mentor a team of Forward Deployed Engineers; design, implement, and deploy Baseten solutions end-to-end for customers; optimize and enhance AI/ML projects; turn vague objectives into clear specs and well-defined PoCs.
Seniority
Manager, hands-on IC