Senior DevOps Engineer – AI Platform / AWS / GPU Infrastructure F/H
Core
Design, secure, and scale GPU-intensive cloud and on-prem infrastructure for sovereign generative AI platforms (Mistral AI, Prisme AI) in a critical banking environment.
Role type
Senior DevOps / Platform Engineer (AI Infrastructure)
Builds
Enterprise-scale LLM serving platforms, inference APIs, and GPU workloads for generative AI agents.
Domain
Banking / Financial Services, Generative AI, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes production administration, AWS hybrid cloud architecture, CI/CD pipeline industrialization, GPU resource allocation optimization, Infrastructure as Code, real-time observability stack implementation, LLM orchestration, security compliance for regulated environments.
Preferred skills
Generative AI platform experience, self-hosted LLM deployment, NVIDIA stack and GPU operators, MLOps/LLMOps, banking/finance background, SRE reliability engineering.
Technologies
AWS, Kubernetes, Docker, Helm, Kustomize, GitLab CI, GitHub Actions, ArgoCD, Terraform, Ansible, Prometheus, Grafana, ELK, Loki, OpenTelemetry, Mistral AI, Prisme AI.
Responsibilities
Design and maintain highly available cloud and on-prem infrastructure for generative AI platforms; Deploy and administer dedicated Kubernetes clusters for AI/LLM workloads; Optimize CPU/GPU/memory/storage allocation and horizontal/vertical scaling; Build and industrialize CI/CD pipelines for AI models and agent applications; Administer complex distributed container environments with high-performance scheduling and autoscaling; Implement advanced observability stacks with AI-specific metrics (inference latency, token generation rate, GPU utilization).