Infrastructure Operations Engineer (APAC)
Core
Design, build, and operate large-scale GPU infrastructure platforms to ensure reliability, automation, and efficient customer provisioning for AI systems.
Role type
Senior Infrastructure Operations Engineer (GPU/Cloud)
Builds
Next-generation AI infrastructure platform, automation systems, and operational workflows for GPU environments.
Domain
Cloud Infrastructure, GPU Computing, Data Center Operations
Deliverable
production ML models | infrastructure
Required skills
Linux server administration, AWS, Kubernetes, Terraform, Ansible, network storage management (NFS, Ceph), monitoring (Prometheus, ELK), Python/Go/bash scripting, networking fundamentals
Preferred skills
Bare metal hardware troubleshooting, GPU server management, datacenter networking (400Gb ethernet, Infiniband), VAST storage systems, SONiC switches, Palo Alto firewalls, Juniper Networks
Responsibilities
Design and roll out new platforms and patterns to minimize incidents; Deploy updates for internal and customer use cases; Collaborate with Infrastructure Engineering, Network Operations, and Software teams; Participate in on-call rotation for incident response.
Seniority
Senior, hands-on IC