LLM Inference Frameworks and Optimization Engineer
Core
Design, develop, and optimize distributed inference engines for large language models (LLMs) and multimodal/vision models to ensure low-latency, high-throughput, and cost-efficient deployment.
Role type
Senior IC inference frameworks and optimization engineer
Builds
Distributed inference engines and serving pipelines for text, image, and multimodal generation models
Domain
AI Infrastructure / Deep Learning Systems
Deliverable
production ML models
Required skills
Distributed systems design, GPU programming (CUDA/Triton), LLM inference frameworks (TensorRT-LLM, vLLM, SGLang, TGI), KV cache systems (Mooncake, PagedAttention), model quantization, compiler optimization, Python, C++
Preferred skills
RDMA/RoCE networking, distributed filesystems (3FS, HDFS, Ceph), Kubernetes orchestration, open-source contributions
Technologies
CUDA, TensorRT, PyTorch, torch.compile, Triton, Kubernetes, RDMA
Responsibilities
Design fault-tolerant, high-concurrency distributed inference engines; Implement parallelism strategies (MoE, tensor, pipeline); Optimize inference using CUDA graphs and speculative decoding; Collaborate on software-hardware co-design for GPUs/TPUs/custom accelerators; Develop efficient model execution plans and E2E serving pipelines
Seniority
Senior, hands-on IC