CareerPlanSign in

大模型应用研发工程师(推理部署优化方向)-TRAE

深圳💼 Full-time🗓 2026-09-28

Core

Design and iterate model inference and deployment solutions for an AI coding product (TRAE) to ensure stability, optimize end-to-end performance, and reduce costs.

Role type

Senior IC LLM inference and deployment optimization engineer

Builds

AI coding agent product (TRAE) serving To C/To B users

Domain

Generative AI / LLM / Cloud Infrastructure

Deliverable

production ML models

Required skills

LLM deployment, vLLM, TRT-LLM, SGLang, CUDA kernel development, GPU hardware optimization, model quantization, MoE sparse structures, Diffusion models

Preferred skills

End-to-end performance analysis, system stability troubleshooting, proactive learning of LLM architectures

Technologies

vLLM, TRT-LLM, SGLang, CUDA, NVIDIA GPUs

Responsibilities

Handle online alerts and manage deployment scaling for model services; Analyze end-to-end latency and throughput to optimize code completion and Agent performance; Design and implement inference pipelines including quantization and acceleration for new model structures.

Sourced via bytedance · Listed on CareerPlan, which tracks 844,000+ jobs from 20+ sources.