AI系统性能工程师-国际化内容安全平台
Core
Design and develop a next-generation multimodal inference architecture and global scheduling system for massive heterogeneous compute resources to handle high-concurrency traffic.
Role type
Senior IC AI Infrastructure Engineer (Multimodal Inference)
Builds
High-throughput, low-latency distributed inference engines for 200B+ parameter models and Visual Language Models (VLMs).
Domain
AI Infrastructure / Large Language Models / Multimodal Systems
Deliverable
production ML models
Required skills
Distributed scheduling systems, High-performance computing, CUDA/Triton operator optimization, GPU microarchitecture, vLLM/SGLang framework internals, Asynchronous scheduling, Resource pooling, Load balancing, MoE architectures, KV Cache optimization, PTQ/QAT, Multimodal model engineering.
Preferred skills
AI Agent for system tuning, Kernel optimization, Production environment customization.
Responsibilities
Design global scheduling systems for heterogeneous compute; Optimize distributed inference strategies (TP/EP/DP); Develop high-performance low-level operators; Build AI-driven infrastructure for model operations and maintenance.
Seniority
Senior, hands-on IC