Technical Operations Lead - AI Cloud Platform
Core
Oversee day-to-day operations, reliability, and support of a large-scale private AI cloud platform.
Role type
Senior Technical Operations Lead (AI Cloud Platform)
Builds
Large-scale private AI cloud platform
Domain
Cloud Infrastructure & AI Operations
Deliverable
infrastructure
Required skills
Linux, Kubernetes, private cloud operations, incident management, disaster recovery, capacity planning, root cause analysis, vendor coordination
Preferred skills
AI/GPU platform operations, SRE practices, ITIL, ISO 27001/22301 familiarity
Technologies
Canonical OpenStack, Canonical Kubernetes, MAAS, Juju, Ceph, Ubuntu Server, Prometheus, Grafana, Elasticsearch, OpenSearch, Loki, Terraform, Ansible, Azure DevOps, Keycloak, Vault, PAM, SIEM, NVIDIA GPU infrastructure
Responsibilities
Lead platform operations and technical support; Ensure platform availability, performance, and service reliability; Own incident, problem, and change management processes; Coordinate major incident response and technical escalations; Maintain operational procedures, runbooks, and support models; Lead backup, restore, and disaster recovery testing; Monitor platform capacity, availability, and performance; Coordinate maintenance activities, upgrades, and vendor support; Drive root cause analysis and continuous improvement initiatives; Support operational readiness for new services and capabilities
Seniority
Senior, hands-on IC with leadership responsibilities
