AI Inference Engineer
Core
Optimizing Large Language Models (LLMs) for inference across diverse environments from data centers to edge devices, focusing on maximizing throughput and minimizing latency.
Role type
Senior IC AI Inference Engineer
Builds
High-performance inference engines and scalable AI serving infrastructure
Domain
AI/ML Infrastructure, High-Performance Computing, Cloud Infrastructure
Deliverable
production ML models
Required skills
Python, C++, Rust, Golang, vLLM, TensorRT, Llama.cpp, Ollama, Docker, Kubernetes, AWS, GCP, Azure, NVIDIA GPU optimization, TPU optimization
Preferred skills
Speculative Decoding, PagedAttention, open-source inference library contributions, CUDA kernel development, MLOps, SRE
Technologies
vLLM, TGI, NVIDIA Triton, TensorRT, Llama.cpp, Ollama, CUDA, CoreML, Kubernetes, Docker
Responsibilities
Build and maintain robust inference engines; Profile and optimize models for specialized hardware backends; Design and implement auto-scaling architectures for inference pipelines; Establish robust observability frameworks for performance monitoring.
Seniority
Senior, hands-on IC