大模型推理工程师(J101025)
Core
Deploy, optimize, and serve large language and multimodal models to support a stable, low-cost MaaS inference platform.
Role type
Senior IC large model inference engineer
Builds
High-throughput, low-latency inference services for MaaS platforms
Domain
AI/ML inference, cloud infrastructure, GPU resource management
Deliverable
production ML models
Required skills
LLM inference optimization, vLLM/TGI/TensorRT, quantization, KV Cache, PagedAttention, dynamic batching, Python, Linux, asynchronous programming, containerization, Kubernetes, high-concurrency service design
Preferred skills
MaaS platform experience, cloud model services, AI middle platform experience, GPU resource scheduling
Responsibilities
Optimize inference latency and throughput, build scalable service architectures, troubleshoot production issues, manage model versioning and deployment, participate in resource management and billing systems
Seniority
Mid-level, hands-on IC