CareerPlanSign in

Inference Systems Backend Engineer

USA💼 Full-time🗓 2026-09-07 → 2026-09-25

Core

Develop and performance-optimize large model training and inference systems, including distributed LLM inference and thousand-GPU training clusters.

Role type

Senior IC backend engineer (large-scale ML inference/training systems)

Builds

Distributed LLM inference systems, high-concurrency training clusters, and elastic GPU scheduling platforms

Domain

Cloud infrastructure, large-scale machine learning, distributed systems

Deliverable

production ML models | infrastructure

Required skills

C/C++, Python, Linux, distributed systems, machine learning frameworks (TensorFlow, PyTorch, MXNet), large model inference frameworks (vLLM, TensorRT-LLM, SGLang, Megatron-LM), subgraph matching, compiler optimization, model quantization, GPU/NPU/TPU integration, elastic scheduling, GPU oversubscription, task orchestration

Preferred skills

large-scale distributed-system architecture design, GPU hardware architecture and software stacks (CUDA, cuDNN), GPU performance analysis and optimization, research background in distributed systems, parallel computing, programming languages, compilers, networking, or storage systems

Technologies

vLLM, TensorRT-LLM, SGLang, Megatron-LM, CUDA, cuDNN, TensorFlow, PyTorch, MXNet

Responsibilities

Optimize model computation and tune thousand-GPU training clusters; build distributed LLM inference systems; schedule large-scale inference traffic; research and introduce architectures involving subgraph matching, compiler optimization, and model quantization; integrate GPUs, NPUs, and TPUs with training and inference frameworks; improve utilization across globally distributed GPU clusters through elastic scheduling, GPU oversubscription, and task orchestration; collaborate with algorithm teams to jointly optimize algorithms and systems

Seniority

Senior, hands-on IC

Sourced via codingjobboard · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.