AI Infrastructure Engineer
Core
Operate and maintain high-performance AI infrastructure orchestrating thousands of GPUs across multiple data centers to support large-scale autonomous driving model training.
Role type
Senior IC AI Infrastructure Engineer
Builds
Scalable, high-availability GPU clusters and distributed training environments for autonomous driving models
Domain
AI Infrastructure / High-Performance Computing / Autonomous Driving
Deliverable
infrastructure
Required skills
Linux kernel operations, Kubernetes orchestration, Slurm workload management, Python scripting, Shell scripting, GPU hardware troubleshooting, TCP/IP networking fundamentals, distributed system architecture design
Preferred skills
Observability stack implementation (Prometheus, Grafana, Datadog), Public cloud infrastructure (AWS, GCP), NVIDIA accelerated computing stack (CUDA, NCCL), Deep learning frameworks (PyTorch, TensorFlow), Infrastructure as Code (Terraform)
Technologies
Kubernetes, Slurm, Docker, Python, Shell, Prometheus, Grafana, Datadog, AWS, GCP, CUDA, NCCL, Terraform, PyTorch, TensorFlow
Responsibilities
Operate and maintain large-scale GPU clusters using Kubernetes and Slurm; Monitor and diagnose failures across GPU hardware and software stacks; Develop automation tools and scripts to streamline infrastructure management; Manage GPU resource quotas and provide technical support to ML researchers; Participate in architectural design and performance tuning of distributed training environments
Seniority
Senior, hands-on IC