Lead Solution Architect
Core
Design and scope enterprise infrastructure solutions for AI/ML and HPC workloads, integrating container orchestration, GPU-accelerated platforms, and high-performance computing clusters.
Role type
Senior Solution Architect (Infrastructure & AI/HPC)
Builds
Production-grade Kubernetes platforms, HPC clusters, and AI/ML infrastructure for enterprise customers.
Domain
Enterprise IT Infrastructure, High-Performance Computing (HPC), Artificial Intelligence/ML
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (OpenShift, Rancher, CNCF), GPU-accelerated workloads (NVIDIA GPU Operator, DCGM), HPC cluster architecture (Lustre, WEKA, InfiniBand, Slurm), Linux system administration (RHEL, SLES, Ubuntu), virtualization (KVM, OpenShift Virtualization), AI platform enablement (NVIDIA AI Enterprise), DevOps/MLOps (CI/CD, IaC)
Preferred skills
Experience with Dell, Lenovo, Supermicro, NVIDIA DGX/HGX platforms, cloud-native security
Technologies
Red Hat OpenShift, SUSE Rancher, NVIDIA GPU Operator, DCGM, Lustre, WEKA, InfiniBand, Mellanox, RoCE, HPE Cluster Management, NVIDIA Base Command Manager, Slurm, Altair PBS Pro, KVM, NVIDIA AI Enterprise
Responsibilities
Architect HPC clusters with GPU/compute nodes and storage technologies; integrate Kubernetes with AI/ML workloads in hybrid/private cloud; deploy and support NVIDIA AI Enterprise and related frameworks; deliver infrastructure projects for AI/ML and HPC; align technical solutions with business goals.
Seniority
Senior, hands-on IC