CareerPlanSign in

Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking

San Francisco🌐 Remote💼 Full-time🗓 2026-08-04 → 2026-09-26

Core

Hands-on technical partner for strategic customers, specializing in inference optimization, fine-tuning pipelines, and production deployment of high-quality AI models.

Role type

Senior IC Forward Deployed Engineer (Inference & Post-Training)

Builds

Optimized inference endpoints and production-ready fine-tuning pipelines for enterprise AI teams

Domain

Generative AI / Large Language Models (LLMs) / Cloud Infrastructure

Deliverable

production ML models

Required skills

Inference engine optimization, KV cache tuning, speculative decoding, tensor parallelism, quantization strategies, LoRA, SFT, DPO, RLHF, GRPO, Python, open-source LLM deployment

Preferred skills

RL training system design, model selection judgment, production environment experience

Technologies

vLLM, TensorRT-LLM, SGLang

Responsibilities

Select and optimize inference engines based on hardware and workload; tune configurations for throughput and latency targets; drive hands-on RL training runs and optimize system design; act as primary technical contact for strategic accounts; establish direct alignment with customers at onboarding; influence software and model roadmap with field insights

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.