Infrastructure Operations Engineer
Core
Day-to-day operations, incident resolution, and service transition for a large-scale private AI cloud platform.
Role type
Senior Infrastructure Operations Engineer
Builds
Availability and stability of critical AI infrastructure (compute, storage, networking, cloud)
Domain
Cloud Infrastructure & AI Platform Operations
Deliverable
infrastructure
Required skills
Linux system administration, Kubernetes, incident troubleshooting, change management, automation scripting, capacity monitoring, disaster recovery
Preferred skills
ITIL practices, Bash/Python scripting, 24x7 operations experience, regulated environment experience
Technologies
Ubuntu Server, Kubernetes, Prometheus, Grafana, Ceph, OpenStack, MAAS, Juju, Terraform, Ansible, Azure DevOps, Elasticsearch, OpenSearch, Loki, NVIDIA GPU
Responsibilities
Investigate and resolve infrastructure/platform/data center incidents; Perform day-to-day operational activities across compute, storage, networking, and cloud; Execute operational procedures, maintenance, and platform upgrades; Monitor platform health, capacity, availability, and performance; Support backup, restore, and disaster recovery activities; Maintain operational documentation, runbooks, and SOPs; Contribute to automation initiatives and continuous operational improvement; Participate in knowledge transfer and service transition activities.
Seniority
Mid-Senior, hands-on IC
