Inference Engineer
Core
Design and build low latency, scalable, and reliable model inference and serving stack for cutting-edge foundation models including Transformers, SSMs, and hybrid models.
Role type
Senior Inference Engineer (Systems & ML)
Builds
Real-time multimodal intelligence products and inference infrastructure
Domain
Generative AI, Large Language Models, State Space Models
Deliverable
production ML models
Required skills
Distributed systems engineering, Inference pipeline design, Machine Learning model implementation, Generative AI experience, CUDA programming, Triton, vLLM, SGLang, Continuous Batching
Preferred skills
Experience with hybrid models, Zero-to-one execution
Technologies
Transformers, SSMs, vLLM, SGLang, CUDA, Triton
Responsibilities
Design and build low latency, scalable, and reliable model inference and serving stack; Work closely with research and product engineers to serve products in a fast, cost-effective, and reliable manner; Design and build robust inference infrastructure and monitoring for products.
Seniority
Senior, hands-on IC