Senior Solutions Architect, Cloud Infrastructure and DevOps - NVIS
Core
Design, build, and maintain large-scale HPC/AI clusters and networking infrastructure for academic and commercial customers.
Role type
Senior Solutions Architect, Cloud Infrastructure and DevOps
Builds
Large-scale AI/HPC systems, networking projects, and automated infrastructure environments
Domain
High-Performance Computing (HPC), Artificial Intelligence (AI), Cloud Infrastructure, Networking
Deliverable
production ML models | infrastructure
Required skills
Linux internals and administration, Kubernetes orchestration, HPC cluster management, TCP/IP and data center architecture, Python scripting, CI/CD pipeline development, storage solutions (Lustre, GPFS, ZFS, XFS), automation tools (Ansible, Jenkins, Puppet/Chef), troubleshooting at bare metal to application level
Preferred skills
GPU architecture knowledge, CUDA/DGX experience, RDMA fabrics (InfiniBand, RoCE), microservice technologies
Technologies
Kubernetes, Slurm, Singularity, Redhat/CentOS, Ubuntu, Lustre, GPFS, ZFS, XFS, InfiniBand, RoCE, Jenkins, Ansible, Puppet, Chef, CUDA, DGX
Responsibilities
Maintain large scale HPC/AI clusters with monitoring, logging, and alerting; Develop and maintain continuous integration and delivery pipelines; Develop tooling to automate deployment and management of large-scale infrastructure; Deploy monitoring solutions for servers, network, and storage; Perform troubleshooting from bare metal to application level; Document standard methodologies and support R&D activities and POCs
Seniority
Senior, hands-on IC