Staff/Principal AI Transformation Engineer
Core
Design and integrate frontier LLM ecosystems and commercial AI platforms to build a unified distributed AI platform for data processing, model training, inference, and evaluation.
Role type
Staff/Principal AI Transformation Engineer (Systems & Infrastructure)
Builds
Unified distributed AI platform, AI DevOps, AI Copilots, resilient automation pipelines, and large-scale production infrastructure.
Domain
Autonomous driving, AI infrastructure, distributed systems, large-scale compute clusters.
Deliverable
production ML models | infrastructure
Required skills
System-level architecture design, C++, Python, Java, JavaScript, PyTorch, Ray, vLLM, Triton Inference Server, Kubernetes, DeepSpeed, Megatron-LM, distributed LLM training, inference optimization, GPU cluster management, hands-on core code development and debugging.
Preferred skills
Autonomous driving algorithms, robotics, physics-based simulation engines, AI Copilot applications, autonomous multi-agent frameworks, open-source maintenance, technical authorship.
Technologies
PyTorch, Ray, vLLM, Triton Inference Server, Kubernetes, DeepSpeed, Megatron-LM, C++, Python, Java, JavaScript.
Responsibilities
Lead evaluation and integration of frontier LLM ecosystems; own architecture of unified distributed AI platform; embed with autonomous driving teams to deliver production-ready AI ecosystems; design automation pipelines for LLM deployment and monitoring; optimize GPU cluster utilization and inference latency; define technical roadmap for AI infrastructure; write core code and debug deep system issues.
Seniority
Staff/Principal, hands-on IC with strategic influence