Software Engineer — Distributed LLM Inference Systems
Core
Design, develop, and optimize distributed inference systems for large language models (LLMs) across diverse hardware architectures.
Role type
Software Engineer (Distributed LLM Inference Systems)
Builds
Distributed inference algorithms, model execution components, and communication layers for AI frameworks.
Domain
Artificial Intelligence / Deep Learning / Distributed Systems
Deliverable
production ML models
Required skills
Python, C++, deep learning frameworks (PyTorch), distributed algorithms, performance debugging, machine learning fundamentals
Preferred skills
Distributed LLM inference/serving, open-source contributions, LLM inference concepts (prefill/decode, KV cache, continuous batching), inference engines (vLLM, SGLang, TensorRT-LLM), AI Agent architecture
Technologies
PyTorch, vLLM, SGLang, TensorRT-LLM
Responsibilities
Implement distributed algorithms (model/data parallel, async communication), develop request schedulers and KV cache management, profile inference workloads for bottlenecks, collaborate to improve latency/throughput/scalability, contribute code/tests/docs to internal and open-source projects
Seniority
Junior (0-1 years experience)