AI Infrastructure Engineer III
Core
Design, operate, and automate scalable AI/ML infrastructure and GPU clusters to enable data scientists and ML engineers to develop, train, and deploy models.
Role type
Senior IC AI Infrastructure Engineer
Builds
Self-service AI platforms, GPU clusters, model serving infrastructure, and MLOps tooling
Domain
Cloud-native AI/ML infrastructure and GPU computing
Deliverable
production ML models
Required skills
Kubernetes, Kubeflow, MLflow, GPU infrastructure management, distributed training, model serving, Infrastructure as Code, Python/Bash/Go scripting, observability
Preferred skills
PyTorch, TensorFlow, Hugging Face, Ray, DeepSpeed, Vector Databases, LLM infrastructure, RAG architectures
Technologies
Kubeflow, MLflow, KServe, Ray, NVIDIA GPU Operator, Terraform, Helm, GitOps, Ansible, Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, AWS, GCP, OCI, Azure
Responsibilities
Design and deploy enterprise AI/ML platforms; Build self-service platforms for Data Scientists and ML Engineers; Design and operate GPU clusters for large-scale AI workloads; Build CI/CD pipelines for ML workloads; Automate AI infrastructure provisioning using Infrastructure as Code; Implement monitoring and observability for GPU utilization and model serving
Seniority
Senior, hands-on IC