Principal Engineer - AI Platform & Operations
Core
Architect and own the technical vision for foundational AI platform systems enabling model serving, feature stores, and experiment tracking for all AI product teams.
Role type
Principal Engineer (AI Platform & Operations)
Builds
Self-serve internal developer platforms for deploying, monitoring, and scaling AI models
Domain
Business travel technology / AI Infrastructure / MLOps
Deliverable
production ML models
Required skills
ML infrastructure design, distributed systems, Kubernetes, cloud ecosystems (AWS/GCP/Azure), Python, LLM inference optimization, platform architecture, technical mentorship
Preferred skills
Experience with Triton, vLLM, Ray Serve, MLflow, Weights & Biases, Kubeflow, Docker, Helm
Technologies
Triton, vLLM, Ray Serve, Kubernetes, AWS, GCP, Azure, Python, MLflow, Weights & Biases, Kubeflow, Docker, Helm
Responsibilities
Define long-term technical roadmap for AI platform covering model serving and CI/CD; Establish engineering benchmarks for deployment and A/B testing; Lead strategies for GPU/compute efficiency and LLM inference optimization; Design monitoring and alerting systems for AI workloads; Drive platform stability and partner with application teams; Mentor senior engineers and lead architecture reviews
Seniority
Principal, hands-on IC with strategic ownership