CareerPlanSign in

Software Engineer — Distributed LLM Inference Systems

PRC, Shanghai💼 Full-time🗓 2026-08-14 → 2026-09-25

Core

Design, develop, and optimize distributed inference systems for large language models (LLMs) across diverse hardware architectures.

Role type

Software Engineer (Distributed LLM Inference Systems)

Builds

Distributed inference algorithms, model execution components, and communication layers for AI frameworks.

Domain

Artificial Intelligence / Deep Learning / Distributed Systems

Deliverable

production ML models

Required skills

Python, C++, deep learning frameworks (PyTorch), distributed algorithms, performance debugging, machine learning fundamentals

Preferred skills

Distributed LLM inference/serving, open-source contributions, LLM inference concepts (prefill/decode, KV cache, continuous batching), inference engines (vLLM, SGLang, TensorRT-LLM), AI Agent architecture

Technologies

PyTorch, vLLM, SGLang, TensorRT-LLM

Responsibilities

Implement distributed algorithms (model/data parallel, async communication), develop request schedulers and KV cache management, profile inference workloads for bottlenecks, collaborate to improve latency/throughput/scalability, contribute code/tests/docs to internal and open-source projects

Seniority

Junior (0-1 years experience)

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.