CareerPlanGet AI match score →

LLM Inference Frameworks and Optimization Engineer

💼 Full-time💰 $160,000–$160,000🗓 2026-07-10 → 2026-07-31

Core

Design, develop, and optimize distributed inference engines for large language models (LLMs) and multimodal/vision models to ensure low-latency, high-throughput, and cost-efficient deployment.

Role type

Senior IC inference frameworks and optimization engineer

Builds

Distributed inference engines and serving pipelines for text, image, and multimodal generation models

Domain

AI Infrastructure / Deep Learning Systems

Deliverable

production ML models

Required skills

Distributed systems design, GPU programming (CUDA/Triton), LLM inference frameworks (TensorRT-LLM, vLLM, SGLang, TGI), KV cache systems (Mooncake, PagedAttention), model quantization, compiler optimization, Python, C++

Preferred skills

RDMA/RoCE networking, distributed filesystems (3FS, HDFS, Ceph), Kubernetes orchestration, open-source contributions

Technologies

CUDA, TensorRT, PyTorch, torch.compile, Triton, Kubernetes, RDMA

Responsibilities

Design fault-tolerant, high-concurrency distributed inference engines; Implement parallelism strategies (MoE, tensor, pipeline); Optimize inference using CUDA graphs and speculative decoding; Collaborate on software-hardware co-design for GPUs/TPUs/custom accelerators; Develop efficient model execution plans and E2E serving pipelines

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗