Principal AI Cloud Engineer
Core
Design, build, and operate cloud-native infrastructure for machine learning and generative AI workloads, including compute, networking, storage, databases, and security.
Role type
Principal AI Cloud Engineer
Builds
Cloud-native AI solutions, model serving infrastructure, and LLMOps platforms for a global financial group
Domain
Financial Services / Cloud Infrastructure / AI/ML
Deliverable
production ML models
Required skills
Large-scale distributed cloud infrastructure, Azure/AWS, GPU clusters, Kubernetes, CI/CD, observability, Infrastructure as Code (Terraform/Bicep), networking, security, cloud-native patterns, MLOps/LLMOps, Python, Go/TypeScript
Preferred skills
GPU optimization (CUDA/NCCL/TensorRT-LLM), Prometheus/Grafana/OpenTelemetry, event streaming (Kafka/Azure Event Hubs), AI platform products (Azure ML/MLflow/KServe/Hugging Face)
Technologies
AWS, Azure, Kubernetes, Terraform, Bicep, Python, Go, TypeScript, Kafka, Prometheus, Grafana, OpenTelemetry, MLflow, KServe, vLLM, RAG, CUDA, NCCL, TensorRT-LLM
Responsibilities
Design and operate cloud-native infrastructure for AI/LLM workloads; establish monitoring and reliability practices; build CI/CD and GitOps pipelines; manage AI infrastructure costs; enable secure AI platform services; support RAG systems; define cloud AI architecture; implement security and Responsible AI controls; lead infrastructure discovery and solution design; operate platforms using SRE practices; mentor engineers; define reusable infrastructure modules; operationalize LLMs with controls; guide platform roadmaps.
Seniority
Principal, strategy & mentorship
