CareerPlanSign in

Senior AI Engineer

Vietnam💼 Full-time🗓 2026-09-09 → 2026-09-26

Core

Design and build local LLM serving environments on GPU hardware to turn AI models into fast, cost-effective production-grade services.

Role type

Senior IC machine-learning infrastructure engineer (LLM inference)

Builds

Production-grade LLM inference services optimized for GPU hardware

Domain

Cybersecurity / AI Infrastructure

Deliverable

production ML models

Required skills

LLM inference optimization, model compression (quantization, pruning, distillation), GPU architecture, CUDA, high-throughput serving engines (vLLM, TensorRT-LLM, TGI), Python, mixed precision strategies

Preferred skills

Custom CUDA/Triton kernel tuning, multi-GPU distributed inference, CI/CD for cloud GPU deployment

Technologies

NVIDIA CUDA, cuDNN, vLLM, TensorRT-LLM, TGI, GPTQ, AWQ, SmoothQuant, FlashAttention, FP8/FP4/INT8/INT4

Responsibilities

Design and build local LLM serving environments on GPU hardware; Optimize LLM models for efficient GPU serving using quantization and compression techniques; Deploy and tune high-throughput serving engines; Establish quality regression gates and run A/B tests for quantized models; Explore and implement novel inference optimization techniques

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.