Infrastructure Operations Engineer
Core
Scale and operate next-generation AI infrastructure platforms, focusing on GPU environments, Linux systems, and automation to ensure reliability and efficiency.
Role type
Senior Infrastructure Operations Engineer
Builds
Large-scale GPU infrastructure, automation systems, and operational workflows for AI training and inference.
Domain
AI Infrastructure / Cloud Computing / Data Center Operations
Deliverable
infrastructure
Required skills
Linux administration, AWS, Kubernetes, Terraform, Ansible, network storage management, monitoring systems, Python/Go/bash scripting, networking fundamentals
Preferred skills
Bare metal hardware troubleshooting, GPU server management, network switch/router/firewall expertise, VAST storage systems
Technologies
AWS, Kubernetes, Terraform, Ansible, Prometheus, ELK stack, GitOps, Python, Go, bash, NFS, Ceph, VAST, SONiC, Palo Alto, Juniper
Responsibilities
Design and roll out new platforms and patterns to minimize incidents; Deploy updates for internal and customer use cases; Collaborate with Infrastructure Engineering, Network Operations, and Software teams; Participate in on-call rotation for incident response.
Seniority
Senior, hands-on IC