AI异构硬件推理优化专家 - Seed Model
Core
Optimizing inference performance for ByteDance's self-developed LLMs on large-scale heterogeneous AI accelerator clusters.
Role type
Senior IC AI inference optimization engineer (heterogeneous hardware)
Builds
Distributed inference frameworks, high-performance operators, and online serving stability for Doubao and Volcengine.
Domain
AI infrastructure / HPC / Distributed systems
Deliverable
production ML models
Required skills
C/C++, Python, Linux, computer architecture, chip microarchitecture, high-performance computing, distributed systems, parallel computing
Preferred skills
AI chip performance optimization, inference serving stacks (vLLM, SGLang, PagedAttention), CUDA, AscendC, TileLang, Triton, CUTLASS, TVM, MLIR, TorchInductor
Responsibilities
Deploy and tune self-developed LLMs on large-scale AI accelerator clusters; optimize scheduling, Batching, KV Cache, memory management, distributed parallelism, load balancing, speculative inference, sparse computation, and quantization; develop and optimize key LLM operators (Attention, GEMM, quantization, compute-compute fusion) for different heterogeneous hardware ISAs.
Seniority
Senior, hands-on IC