Engineer, Inference & Model serving
Core
Building high-performance, low-latency inference systems for LLMs, speech, and vision models in production.
Role type
Senior IC ML Model Serving Engineer
Builds
Real-time AI systems serving LLMs, speech, and vision models
Domain
Artificial Intelligence / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
ML inference or model serving systems, Python, PyTorch, distributed systems, production infrastructure, latency and throughput optimisation, GPU profiling, CUDA, Kubernetes, Ray
Preferred skills
Systems or performance engineering mindset
Technologies
vLLM, TensorRT-LLM, Triton, SGLang, CUDA, Kubernetes, Ray
Responsibilities
Building high-performance serving systems for LLM, speech, and vision models, Scaling inference to production workloads with strict latency requirements, Optimising GPU utilisation and execution efficiency, Implementing techniques like continuous batching, KV cache optimisation, speculative decoding, and prefill/decode separation, Improving frameworks such as vLLM, TensorRT-LLM, Triton, and SGLang, Profiling and debugging performance across GPU, memory, and system layers
Seniority
Senior, hands-on IC