大模型应用研发工程师(推理部署优化方向)-TRAE
Core
Design and iterate model inference and deployment solutions for an AI coding product (TRAE) to ensure stability, optimize end-to-end performance, and reduce costs.
Role type
Senior IC LLM inference and deployment optimization engineer
Builds
AI coding agent product (TRAE) serving To C/To B users
Domain
Generative AI / LLM / Cloud Infrastructure
Deliverable
production ML models
Required skills
LLM deployment, vLLM, TRT-LLM, SGLang, CUDA kernel development, GPU hardware optimization, model quantization, MoE sparse structures, Diffusion models
Preferred skills
End-to-end performance analysis, system stability troubleshooting, proactive learning of LLM architectures
Technologies
vLLM, TRT-LLM, SGLang, CUDA, NVIDIA GPUs
Responsibilities
Handle online alerts and manage deployment scaling for model services; Analyze end-to-end latency and throughput to optimize code completion and Agent performance; Design and implement inference pipelines including quantization and acceleration for new model structures.