Senior Staff Deployment Automation Engineer
Core
Own deployment and integration testing automation for large-scale, multi-node GPU clusters and CI/CD infrastructure for a virtualized AI Cloud.
Role type
Senior Staff/Principal Deployment Automation Engineer
Builds
CI/CD platforms, automated test suites, and control applications for bare-metal Linux systems and virtualized GPU/CPU hosts.
Domain
AI Infrastructure / Data Center Operations / Cloud Computing
Deliverable
production ML models | infrastructure
Required skills
CI/CD pipeline design, Python, Bash, Linux kernel internals, Kubernetes, Terraform, Ansible, multi-node cluster validation, RDMA/RoCE/InfiniBand protocols, NVIDIA CUDA/NCCL or AMD ROCm/RCCL stacks
Preferred skills
MNNVL experience, hardware-level debugging tools (NVIDIA Nsight, AMD Omniperf), containerized GPU orchestration
Technologies
Gitlab, Ansible, AWX, osquery, fio, stress-ng, iperf, Python, Go, Kubernetes, Docker, Terraform, Postgres
Responsibilities
Own deployment and integration testing automation for bare-metal on-premise systems; Build CI/CD platforms for low-level system releases; Design and execute large-scale validation tests for multi-node virtualized clusters; Maintain and scale bare-metal Linux configurations; Create control applications for canary deployments, Blue/Green testing, and automatic rollback; Develop automation frameworks to provision, configure, and stress-test multi-node virtualized environments; Create automated test suites for performance and multi-tenant isolation
Seniority
Senior Staff/Principal, hands-on IC