Machine Learning Engineer - Inference
Core
Design and build production systems to optimize and scale AI inference for large language models.
Role type
Senior IC machine learning engineer (inference systems)
Builds
High-performance inference engine services and runtime systems for large-scale AI applications
Domain
Artificial Intelligence / Large Language Models / Systems Engineering
Deliverable
production ML models
Required skills
Python, PyTorch, high-performance system design, multi-threading, memory management, networking, storage, code reviews, fault-tolerant system implementation
Preferred skills
TGI, vLLM, TensorRT-LLM, Optimum, speculative decoding, CUDA, Triton, Rust, Cython
Technologies
PyTorch, CUDA, Triton, Rust, Cython
Responsibilities
Design and build production systems for the inference engine; Develop and optimize runtime inference services; Collaborate with researchers and engineers to bring new features; Conduct design and code reviews; Create services, tools, and documentation; Implement robust and fault-tolerant systems for data ingestion and processing
Seniority
Senior, hands-on IC