CareerPlanSign in

推理性能优化专家-计算

杭州💼 Full-time🗓 2026-09-28

Core

Lead inference optimization for LLMs and multimodal models, focusing on quantization, sparsity, communication protocols, and core operator efficiency to enable scalable deployment.

Role type

Senior IC inference performance optimization engineer

Builds

Standardized performance benchmarking systems, automated tuning pipelines, and high-throughput distributed inference clusters

Domain

AI/ML infrastructure, distributed systems, high-performance computing

Deliverable

production ML models

Required skills

C++/Python/Go, LLM inference optimization, mixed-precision quantization, sparse attention mechanisms, RDMA/TCP protocol optimization, AI core operator development (GEMM, convolution), Triton/MLIR compilation frameworks, hardware instruction set adaptation (CUDA/ROCm/Ascend/Cambricon), distributed system topology design, speculative decoding, MoE routing strategies

Preferred skills

PyTorch/TensorFlow kernel architecture knowledge, system-level profiling tools (Nsight, PyTorch Profiler), low-latency serialization, traffic scheduling

Technologies

INT4/INT8/FP8, Sparse Attention, RDMA, Triton, MLIR, CUDA, ROCm, Ascend, Cambricon, PyTorch Profiler, NVIDIA Nsight

Responsibilities

Implement hybrid precision quantization and sparsity techniques; optimize cross-node communication stacks and topology; develop and fuse AI core operators for hardware acceleration; design end-to-end scheduling and overlap strategies for MoE architectures

Seniority

Senior, hands-on IC

Sourced via bytedance · Listed on CareerPlan, which tracks 855,000+ jobs from 20+ sources.