CareerPlanSign in

大模型推理优化工程师-抖音

深圳💼 Full-time🗓 2026-09-28

Core

Research and engineering for accelerating inference of AIGC large models (image/video/3D/multimodal) and optimizing GPU resource management during training.

Role type

Senior IC large model inference optimization engineer

Builds

Optimized inference frameworks and deployment services for AIGC models

Domain

AI/ML, Large Language Models, Computer Vision, Graphics

Deliverable

production ML models

Required skills

C++, Python, PyTorch, GPU/NPU operator development, model distillation, quantization, pruning, TensorRT, vLLM, Diffusion model architecture, LLM architecture

Preferred skills

Experience in heterogeneous card adaptation, deep understanding of operator implementation details

Responsibilities

Develop model distillation, quantization, and pruning techniques; adapt models to heterogeneous hardware; optimize GPU utilization during training; deploy engineering services for business use cases; research and implement frontier inference acceleration technologies

Sourced via bytedance · Listed on CareerPlan, which tracks 846,000+ jobs from 20+ sources.