CareerPlanSign in

AML-火山方舟大模型推理系统工程师

上海💼 Full-time🗓 2026-09-28

Core

Design and optimize large-scale training and inference systems for LLMs, managing high-concurrency traffic and heterogeneous hardware clusters.

Role type

Senior IC LLM inference and training systems engineer

Builds

High-performance distributed training clusters and scalable inference systems for LLMs

Domain

Cloud computing / Large Language Models / Distributed Systems

Deliverable

production ML models

Required skills

C/C++, Python, Linux, distributed systems architecture, GPU/NPU/TPU integration, model quantization, subgraph matching, compilation optimization, vLLM, TensorRT-LLM, SGLang, Megatron-LM

Preferred skills

GPU hardware architecture knowledge, CUDA/cuDNN stack expertise, performance analysis, large-scale distributed system design

Responsibilities

Optimize model compute performance and tune thousand-card training clusters; Develop distributed inference systems and manage large-scale traffic scheduling; Research and introduce forward-looking architecture techniques; Integrate heterogeneous hardware into training/inference frameworks; Improve cluster utilization via elastic scheduling and GPU overcommitment; Collaborate with algorithm teams for joint optimization.

Sourced via bytedance · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.