CareerPlanSign in

AI异构硬件推理优化专家 - Seed Model

杭州💼 Full-time🗓 2026-09-28

Core

Optimizing inference performance for ByteDance's self-developed LLMs on large-scale heterogeneous AI accelerator clusters.

Role type

Senior IC AI inference optimization engineer (heterogeneous hardware)

Builds

Distributed inference frameworks, high-performance operators, and online serving stability for Doubao and Volcengine.

Domain

AI infrastructure / HPC / Distributed systems

Deliverable

production ML models

Required skills

C/C++, Python, Linux, computer architecture, chip microarchitecture, high-performance computing, distributed systems, parallel computing

Preferred skills

AI chip performance optimization, inference serving stacks (vLLM, SGLang, PagedAttention), CUDA, AscendC, TileLang, Triton, CUTLASS, TVM, MLIR, TorchInductor

Responsibilities

Deploy and tune self-developed LLMs on large-scale AI accelerator clusters; optimize scheduling, Batching, KV Cache, memory management, distributed parallelism, load balancing, speculative inference, sparse computation, and quantization; develop and optimize key LLM operators (Attention, GEMM, quantization, compute-compute fusion) for different heterogeneous hardware ISAs.

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.