Senior Cloud Infrastructure Engineer
Core
Architect and manage large-scale compute and data infrastructure powering autonomous driving stacks, ensuring multi-GPU clusters and distributed training frameworks are scalable and resilient.
Role type
Senior Cloud Infrastructure Engineer (MLOps)
Builds
High-performance GPU/TPU clusters, data pipelines for sensor processing, and automated model lifecycle workflows for autonomous vehicle simulation and training.
Domain
Autonomous driving / AI infrastructure / Cloud computing
Deliverable
production ML models
Required skills
Kubernetes (K8s) mastery, Infrastructure as Code (Terraform, Helm), Distributed systems (Ray, PyTorch Distributed), Data engineering (Apache Airflow, Kafka, Spark), Observability (Prometheus, Grafana, OpenTelemetry), Python, Bash scripting, IAM/RBAC
Preferred skills
Distributed training frameworks (FSDP, DeepSpeed), AI Agent orchestration (LangGraph, AutoGen), Advanced networking protocols (InfiniBand, RoCE v2), Model Context Protocol (MCP)
Technologies
Kubernetes, NVIDIA GPU Operator, Terraform, Helm, Apache Airflow, Kafka, Spark, ArgoCD, Gitlab CI/CD, MLFlow, Triton Inference Server, Ray Serve, ONNX Runtime, PyTorch, TorchElastic, Horovod, Prometheus, Grafana, OpenTelemetry, LangGraph, CrewAI
Responsibilities
Architect and maintain mission-critical Kubernetes clusters optimized for heavy GPU/TPU workloads; Implement and optimize Kubernetes-native GPU scheduling; Build large-scale data pipelines to process raw sensor data; Develop agent-driven CI/CD workflows for infrastructure and model artifacts; Design and maintain MLFlow and feature store integrations for model tracking; Optimize low-level communication (NCCL, InfiniBand) for large-scale distributed training.
Seniority
Senior, hands-on IC