Data Center Operations System Engineer III (Querétaro, Mexico)
Core
Ensure new server, storage, and network infrastructure is properly racked, labeled, cabled, and configured in advanced GPU and networking systems.
Role type
Senior IC data center operations system engineer
Builds
AI cloud infrastructure for researchers, enterprises, and hyperscalers
Domain
Data center operations, critical infrastructure, AI hardware deployment
Deliverable
production ML models | infrastructure
Required skills
structured cabling, DCIM software, power distribution management, fiber testing, server hardware boot process, cold/hot aisle containment, PDU balancing, Linux administration, ticketing systems (JIRA/Zendesk)
Preferred skills
400Gb Infiniband architectures, DDP/SCM cluster storage systems, High Performance Compute GPU systems (Nvidia NVL72), carrier DIA circuit test and turn ups
Technologies
DCIM software, JIRA, Zendesk, Nvidia NVL72, Infiniband, Linux
Responsibilities
Racking, labeling, and cabling new server/storage/network infrastructure; Troubleshooting hardware and software issues in GPU and networking systems; Documenting and updating data center layout and network topology; Managing parts depot inventory and tracking equipment through deployment; Partnering with HW Support and RMA teams for incident resolution and part replacement; Training junior staff on best practices
Seniority
Senior, hands-on IC