Senior Cloud Infrastructure and DevOps Solutions Architect
Core
Architect and validate large-scale AI/HPC infrastructure solutions for academic and commercial customers, ensuring production stability and performance for GPU clusters.
Role type
Senior Cloud Infrastructure and DevOps Solutions Architect
Builds
Large-scale GPU clusters, Kubernetes-based platforms, and automated HPC/AI environments
Domain
AI/ML, High-Performance Computing (HPC), Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes orchestration, HPC/AI cluster management, Linux systems administration, Python/Bash scripting, Infrastructure-as-Code (Terraform/Ansible), GPU workload profiling, storage systems (Lustre/GPFS/ZFS), observability (Prometheus/Grafana)
Preferred skills
NVIDIA Base Command Manager (BCM), RDMA fabrics (InfiniBand/RoCE), AI-native scheduling frameworks (KAI/Grove/Dynamo), DPU/DOCA infrastructure services
Technologies
Kubernetes, Slurm, KubeVirt, Prometheus, Grafana, Ansible, Terraform, NVIDIA CUDA, InfiniBand, NVLink
Responsibilities
Own full-solution validation including cluster stability testing and burn-in; minimize time to production workload; manage Day 2 production stability and fleet-scale monitoring; assess customer environments and enable third-party ISV workloads; provide consultative guidance and troubleshooting across the full stack; act as technical leader for accounts through knowledge transfer and documentation
Seniority
Senior, hands-on IC with strategic advisory