Infrastructure Engineer
Core
Architect and build secure, multi-tenant observability systems for GPU infrastructure at global scale, enabling AI training and inference workloads.
Role type
Senior Infrastructure Engineer (GPU Observability & Platform Engineering)
Builds
Production-scale monitoring platforms, custom exporters, Kubernetes controllers, and self-service observability interfaces for AI compute providers.
Domain
Enterprise AI Infrastructure, High-Performance Computing (HPC), GPU Cloud
Deliverable
production ML models | infrastructure
Required skills
GPU observability (DCGM, NVML), Systems programming (Go/Rust), Kubernetes controller development, LGTM stack (Loki, Grafana, Tempo, Mimir), SLURM cluster management, GitOps (ArgoCD), Multi-tenant isolation design, SRE practices.
Preferred skills
AI workload monitoring (NCCL, distributed training), Bare-metal systems administration, Cost analytics engineering.
Technologies
DCGM, NVML, Go, Rust, Kubernetes, Prometheus, Loki, Grafana, Tempo, Mimir, Thanos, VictoriaMetrics, SLURM, ArgoCD, OpenTelemetry, systemd.
Responsibilities
Build custom Prometheus exporters for GPU metrics; Develop Kubernetes controllers for GPU workload management; Deploy and manage LGTM stack for production-scale observability; Integrate SLURM clusters with Kubernetes for HPC workloads; Design multi-tenant observability isolation and RBAC; Implement intelligent alerting and SLO tracking for GPU failures and performance degradation.
Seniority
Senior, hands-on IC