CareerPlanSign in

大模型可观测/模型优化工程师(J105368)

北京市💼 Full-time🗓 2026-09-03 → 2026-09-28

Core

Build an observability platform for large language models covering the full lifecycle of training and inference to analyze performance, accuracy, and stability.

Role type

Senior IC AI Infrastructure Engineer (LLM Observability & Optimization)

Builds

AI Infra observability platform, model analysis tools, and optimization toolchains

Domain

AI Infrastructure, Large Language Models, GPU/Heterogeneous Computing

Deliverable

production ML models

Required skills

Python, C++, Linux, PyTorch, CUDA, GPU programming, model quantization, model adaptation, performance profiling, observability platform development

Preferred skills

NPU/XPU experience, training/inference framework knowledge, fault localization

Technologies

PyTorch, CUDA, Linux, GPU, NPU, XPU

Responsibilities

Develop tools for model adaptation, performance optimization, and accuracy issue localization; build metrics, log, trace, and profiling capabilities for GPU/heterogeneous compute environments.

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.