CareerPlanSign in

大模型推理工程师(J101025)

北京市💼 Full-time🗓 2026-07-21 → 2026-09-28

Core

Deploy, optimize, and serve large language and multimodal models to support a stable, low-cost MaaS inference platform.

Role type

Senior IC large model inference engineer

Builds

High-throughput, low-latency inference services for MaaS platforms

Domain

AI/ML inference, cloud infrastructure, GPU resource management

Deliverable

production ML models

Required skills

LLM inference optimization, vLLM/TGI/TensorRT, quantization, KV Cache, PagedAttention, dynamic batching, Python, Linux, asynchronous programming, containerization, Kubernetes, high-concurrency service design

Preferred skills

MaaS platform experience, cloud model services, AI middle platform experience, GPU resource scheduling

Responsibilities

Optimize inference latency and throughput, build scalable service architectures, troubleshoot production issues, manage model versioning and deployment, participate in resource management and billing systems

Seniority

Mid-level, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.