Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking
Core
Hands-on technical partner for strategic customers, specializing in inference optimization, fine-tuning pipelines, and production deployment of high-quality AI models.
Role type
Senior IC Forward Deployed Engineer (Inference & Post-Training)
Builds
Optimized inference endpoints and production-ready fine-tuning pipelines for enterprise AI teams
Domain
Generative AI / Large Language Models (LLMs) / Cloud Infrastructure
Deliverable
production ML models
Required skills
Inference engine optimization, KV cache tuning, speculative decoding, tensor parallelism, quantization strategies, LoRA, SFT, DPO, RLHF, GRPO, Python, open-source LLM deployment
Preferred skills
RL training system design, model selection judgment, production environment experience
Technologies
vLLM, TensorRT-LLM, SGLang
Responsibilities
Select and optimize inference engines based on hardware and workload; tune configurations for throughput and latency targets; drive hands-on RL training runs and optimize system design; act as primary technical contact for strategic accounts; establish direct alignment with customers at onboarding; influence software and model roadmap with field insights
Seniority
Senior, hands-on IC