Devops & SysOps Architect
Core
Senior Lead DevOps & SysOps Architect designing and running mission-critical GPU, HPC, and Kubernetes platforms while executing presales and client delivery.
Role type
Senior Lead IC DevOps & SysOps Architect (GPU/HPC)
Builds
Production-grade multi-cluster Kubernetes platforms, GPU/HPC infrastructure, and client-facing proof-of-concepts
Domain
High-Performance Computing (HPC), AI Infrastructure, Cloud/Hybrid Systems
Deliverable
production ML models | infrastructure
Required skills
Kubernetes (RKE2, EKS, AKS), Linux system administration, GPU infrastructure (NVIDIA H100/A100/B200), Infrastructure as Code (Terraform, Ansible), GitOps (ArgoCD, Fleet), CI/CD (Azure DevOps, Jenkins), Observability (Prometheus, Loki), Security (Vault, IAM, RBAC), HPC storage (Weka, Ceph), InfiniBand/RDMA networking
Preferred skills
NVIDIA GPU Operator, MIG partitioning, Run:AI, TensorRT, Weka, Distributed AI training infrastructure (PyTorch DDP, Horovod, DeepSpeed)
Technologies
RKE2, EKS, AKS, NVIDIA H100, A100, B200, ArgoCD, Fleet, Terraform, Ansible, Azure DevOps, Jenkins, Prometheus, Thanos, Loki, Fluent Bit, DCGM Exporter, NVIDIA Nsight Systems, HashiCorp Vault, Istio, Cilium, Longhorn, Ceph, Weka, InfiniBand, RoCE, CUDA, cuDNN, NCCL, NVIDIA Triton Inference Server, TensorRT
Responsibilities
Partner with sales to identify opportunities and lead technical presales activities including discovery workshops and RFP responses; Operate directly within client accounts to run, troubleshoot, and optimize production Kubernetes clusters and GPU/HPC environments; Design end-to-end platform architectures for cloud, hybrid, and on-premises HPC environments; Implement and maintain IaC pipelines, GitOps workflows, and CI/CD systems; Lead incident response and root cause analysis for critical production issues; Define workload isolation models, networking architectures, and storage strategies for multi-tenant platforms
Seniority
Senior, hands-on IC with strategic presales component