Principal Machine Learning Engineer
Core
Senior technical anchor and solution architect for Grab's core ML and AI infrastructure platform, enabling data scientists and ML engineers to ship models for fraud detection, search ranking, and foundation model efforts.
Role type
Principal Machine Learning Engineer (Platform & Infrastructure)
Builds
Core ML/AI infrastructure including model serving, ML pipelines, data serving, and AI automation tooling for hundreds of internal users.
Domain
Ride-hailing and food delivery superapp; Large-scale distributed ML systems and LLM infrastructure.
Deliverable
production ML models | infrastructure
Required skills
Advanced MLOps & ML Platform Engineering, Distributed Systems & Infrastructure, Architecture Design, AI/LLM System Experience, Innovation & AI Fluency, Adaptive Execution & Ownership, Coaching with Care
Preferred skills
null
Technologies
Kubeflow, MLflow, Triton, TorchServe, PyTorch, Ray, Horovod, vLLM, TensorRT-LLM, Kubernetes, GPUs, TPUs
Responsibilities
Design end-to-end solutions for AIP users and serve as senior technical escalation point; Drive state of large-scale training (throughput, reliability, cost); Optimize end-to-end model iteration loop from idea to shipped model; Design integrations across AIP surfaces for a coherent platform experience; Translate user pain points into functional requirements for AIP teams; Produce reference architectures and best-practice guidance for ML use cases; Define strategic roadmap and mentor senior engineers on SOTA ML infrastructure.
Seniority
Principal, hands-on IC with strategic mentorship