Senior AI Infrastructure Engineer
Core
Design, build, and scale the high-performance AI platform powering autonomous driving models, bridging the gap between research and production to ensure scalability and efficiency.
Role type
Senior AI Infrastructure Engineer
Builds
Distributed training clusters, experiment tracking systems, and production model deployment pipelines for autonomous driving models.
Domain
Autonomous logistics / AI Infrastructure
Deliverable
production ML models
Required skills
Multi-GPU training strategies, Kubernetes, Terraform, PyTorch Distributed, Ray Train, NCCL tuning, Kubernetes-native GPU scheduling, TensorRT, ONNX Runtime, Triton Inference Server, LangGraph, CrewAI, AutoGen, MLFlow, Argo Workflows, Prometheus, Grafana, OpenTelemetry, Apache Airflow, Kafka, Spark, S3, GCS, Delta Lake
Preferred skills
Experience with 3D Gaussian Splatting (3DGS), Agentic AI agents for infrastructure monitoring, A/B testing and shadow deployment strategies
Technologies
PyTorch, Ray, Kubernetes, Terraform, Helm, NCCL, InfiniBand, RoCE v2, TensorRT, ONNX Runtime, Triton, LangGraph, CrewAI, AutoGen, MLFlow, Argo Workflows, Prometheus, Grafana, OpenTelemetry, ELK Stack, Apache Airflow, Kafka, Spark, S3, GCS, Delta Lake
Responsibilities
Scale complex models across multi-node setups using PyTorch Distributed and Ray Train; Architect and optimize multi-GPU setups for efficient model parallelism; Optimize low-level communication (NCCL, InfiniBand) to minimize latency; Deploy and scale optimized model artifacts using TensorRT and Triton; Architect self-healing AI agents to monitor GPU cluster health; Design and maintain ML infrastructure leveraging MLFlow and Argo Workflows; Drive Infrastructure as Code using Terraform and Helm; Define and track key ML system metrics including training convergence and drift detection.
Seniority
Senior, hands-on IC