Datacenter Field Engineer
Core
Own the physical health and foundational infrastructure of GPU clusters, ensuring stable, secure, and fully operational physical environments for demanding compute workloads.
Role type
Senior IC hardware operations & systems engineer (GPU clusters)
Builds
High-density GPU compute clusters for AI research and real-time applications
Domain
AI infrastructure / Datacenter hardware operations
Deliverable
infrastructure
Required skills
Linux systems administration, server hardware troubleshooting, networking security, Bash scripting, vendor coordination
Preferred skills
Configuration management (Ansible/SaltStack/Terraform), high-TDP accelerator cooling management
Technologies
Ubuntu/CentOS/RHEL, NVIDIA H100, AMD MI300, LDAP/Active Directory, NFS/GPFS/Lustre, iptables/firewalld, SSH
Responsibilities
Respond to physical system outages and hardware failures, monitor hardware health (thermals, power, loads), coordinate RMA processes and repairs, rack and cable new GPU nodes, install and patch Linux OS, configure networking and security controls, manage identity and storage systems
Seniority
Senior, hands-on IC