CareerPlanSign in

AI推理工程师(J105801)

深圳市,新加坡共和国💼 Full-time🗓 2026-09-14 → 2026-09-28

Core

Deploy and optimize deep learning model inference across heterogeneous GPU and TPU hardware platforms.

Role type

Senior IC AI inference engineer (GPU/TPU)

Builds

High-throughput, low-latency inference services for AI models

Domain

AI infrastructure / High-performance computing

Deliverable

production ML models

Required skills

GPU/TPU architecture knowledge, CUDA programming, C++, Python, deep learning frameworks (PyTorch/TensorFlow/JAX), model quantization, operator fusion, system programming

Preferred skills

TensorRT, cuDNN, NCCL, Triton Inference Server, XLA, PyTorch/XLA

Technologies

NVIDIA GPU, TPU, CUDA, TensorRT, cuDNN, NCCL, Triton, XLA, PyTorch, TensorFlow, JAX

Responsibilities

Deploy and tune inference performance on GPU and TPU hardware; implement model optimization techniques (quantization, operator fusion); collaborate with algorithm teams on model-inference co-design; build and maintain inference deployment toolchains

Seniority

Senior, hands-on IC

Sourced via baidu · Listed on CareerPlan, which tracks 845,000+ jobs from 20+ sources.