Senior Data Center Deployment Engineer
Core
End-to-end deployment, validation, and production readiness of next-generation GPU platforms inside data centers.
Role type
Senior Data Center Deployment Engineer
Builds
GB-series racks and NVIDIA H200/B200-based AI systems for Nebius's full-stack AI cloud platform
Domain
AI infrastructure / Data Center Hardware / Linux Systems
Deliverable
production ML models
Required skills
Data center infrastructure deployment, GPU-dense systems (NVIDIA H-series), High-density rack deployments (GB-series), Linux troubleshooting and scripting, Hardware/OS/firmware/network diagnostics, Field repair coordination, Technical team leadership, Root cause analysis, Deployment timeline management
Preferred skills
AI/HPC cluster deployment at scale, Automated provisioning systems, Hardware qualification/burn-in testing, Rapid infrastructure expansion, ARM-based/heterogeneous compute environments
Technologies
NVIDIA H200, NVIDIA B200, Linux
Responsibilities
Lead end-to-end deployment of GB-series racks within data center environments, Oversee installation, bring-up, validation, and production readiness of NVIDIA H200 and B200-based servers, Troubleshoot complex hardware, firmware, Linux OS, and networking issues, Execute structured testing and validation procedures during deployment, Develop and maintain basic Linux-based hardware health-check and diagnostic scripts, Coordinate on-site hardware repairs, part replacements, and vendor escalations, Drive root cause analysis and ensure corrective actions are implemented, Manage and prioritize deployment timelines across multiple concurrent rollouts, Provide technical leadership and guidance to on-site engineers and technicians, Partner with networking and infrastructure teams to ensure seamless integration, Document deployment processes, validation standards, and operational runbooks
Seniority
Senior, hands-on IC with leadership responsibilities