AI推理工程师(J105801)
Core
Deploy and optimize deep learning model inference across heterogeneous GPU and TPU hardware platforms.
Role type
Senior IC AI inference engineer (GPU/TPU)
Builds
High-throughput, low-latency inference services for AI models
Domain
AI infrastructure / High-performance computing
Deliverable
production ML models
Required skills
GPU/TPU architecture knowledge, CUDA programming, C++, Python, deep learning frameworks (PyTorch/TensorFlow/JAX), model quantization, operator fusion, system programming
Preferred skills
TensorRT, cuDNN, NCCL, Triton Inference Server, XLA, PyTorch/XLA
Technologies
NVIDIA GPU, TPU, CUDA, TensorRT, cuDNN, NCCL, Triton, XLA, PyTorch, TensorFlow, JAX
Responsibilities
Deploy and tune inference performance on GPU and TPU hardware; implement model optimization techniques (quantization, operator fusion); collaborate with algorithm teams on model-inference co-design; build and maintain inference deployment toolchains
Seniority
Senior, hands-on IC