Senior AI Engineer
Core
Design and build local LLM serving environments on GPU hardware to turn AI models into fast, cost-effective production-grade services.
Role type
Senior IC machine-learning infrastructure engineer (LLM inference)
Builds
Production-grade LLM inference services optimized for GPU hardware
Domain
Cybersecurity / AI Infrastructure
Deliverable
production ML models
Required skills
LLM inference optimization, model compression (quantization, pruning, distillation), GPU architecture, CUDA, high-throughput serving engines (vLLM, TensorRT-LLM, TGI), Python, mixed precision strategies
Preferred skills
Custom CUDA/Triton kernel tuning, multi-GPU distributed inference, CI/CD for cloud GPU deployment
Technologies
NVIDIA CUDA, cuDNN, vLLM, TensorRT-LLM, TGI, GPTQ, AWQ, SmoothQuant, FlashAttention, FP8/FP4/INT8/INT4
Responsibilities
Design and build local LLM serving environments on GPU hardware; Optimize LLM models for efficient GPU serving using quantization and compression techniques; Deploy and tune high-throughput serving engines; Establish quality regression gates and run A/B tests for quantized models; Explore and implement novel inference optimization techniques
Seniority
Senior, hands-on IC