多模态大模型推理服务研发工程师-Data AML(北京/上海/杭州/深圳)
Core
Design and develop online serving systems for SOTA multimodal models (Seedance, Seedream, Seed3D) to support real-time inference for video generation and other multimodal tasks.
Role type
Senior IC machine-learning inference engineer (multimodal)
Builds
High-concurrency, low-latency, high-availability inference services for multimodal and video generation models
Domain
AI / Large Language Models / Multimodal AI
Deliverable
production ML models
Required skills
Go, Python, C++, high-concurrency service architecture, RPC frameworks, asynchronous concurrency, connection pooling, circuit breakers, load balancing, service discovery, message queues, distributed caching, system-level performance profiling, GPU cluster management
Preferred skills
Large model online inference experience, Transformer/DiT/MoE architecture knowledge, inference optimization techniques, inference framework plugin development, model serving toolchain construction
Responsibilities
Design and develop online serving systems for SOTA multimodal models; Participate in inference framework development including request scheduling and model parallelism; Design and implement core service architecture for traffic scheduling and elastic scaling; Optimize end-to-end inference link for GPU utilization and throughput; Lead observability construction and diagnose online issues to ensure stability