AI Platform Engineer, Training and Inference
Core
Building the compute layer for training, evaluating, and serving AI models (LLMs, RL) to power Saviynt's identity security products.
Role type
Senior ML Platform Engineer (Training & Inference)
Builds
Distributed training clusters, multi-engine LLM inference mesh, and full model promotion lifecycle pipelines.
Domain
Identity Security / AI Infrastructure
Deliverable
production ML models
Required skills
Ray ecosystem (Train, Serve, Core, Data), LLM serving engines (vLLM, SGLang, Triton), distributed training (DDP, FSDP, NCCL), RL infrastructure (Flyte, RLlib), model lifecycle operations, Python, PyTorch, vector databases (Pgvector, Qdrant)
Preferred skills
Quantization (GPTQ, AWQ, bitsandbytes)
Technologies
Ray, Kubernetes (GKE), vLLM, SGLang, NVIDIA Triton, Flyte, PyTorch, H100 GPUs, k6, MLflow
Responsibilities
Manage KubeRay on GKE and tune Ray Core Task/Actor scheduling; Operate distributed training with Ray Train on H100 clusters; Build and operate the LLM inference mesh with Ray Serve; Optimize inference performance via fractional GPU allocation and continuous batching; Design and operate the model routing layer; Build RL training infrastructure with Flyte workflows; Operate the full model promotion lifecycle from shadow mode to canary; Operate the retrain pipeline with drift detection and quality gates; Integrate RAG retrieval into the inference mesh.
Seniority
Senior, hands-on IC