Sr. AI Inference Systems Engineer
Core
Lead end-to-end optimization of the full inference pipeline for Large Models (LLM, Multimodal) to maximize throughput and minimize latency.
Role type
Senior IC AI Inference Systems Engineer
Builds
High-performance inference frameworks and optimized inference pipelines for cloud and smart industry solutions
Domain
Cloud computing, AI inference, heterogeneous computing
Deliverable
production ML models
Required skills
AI inference optimization, heterogeneous computing, KV Cache management, Quantization, Intelligent Routing, parallel computing, distributed systems, low-level programming (CUDA, Triton), deep learning frameworks (PyTorch, TensorFlow)
Preferred skills
Ultra-large-scale model optimization, inference cluster tuning, AI inference productization, technical publications or patents
Technologies
CUDA, Triton, PyTorch, TensorFlow
Responsibilities
Design and implement high-performance inference frameworks; optimize scheduling and memory management; conduct research on hardware accelerators; lead efforts to overcome technical bottlenecks; mentor team members
Seniority
Senior, hands-on IC