CareerPlanGet AI match score →

Infrastructure Engineer

Bengaluru, Karnataka, India💼 Full-time🗓 2026-07-14 → 2026-07-31

Core

Architect and build secure, multi-tenant observability systems for GPU infrastructure at global scale, enabling AI training and inference workloads.

Role type

Senior Infrastructure Engineer (GPU Observability & Platform Engineering)

Builds

Production-scale monitoring platforms, custom exporters, Kubernetes controllers, and self-service observability interfaces for AI compute providers.

Domain

Enterprise AI Infrastructure, High-Performance Computing (HPC), GPU Cloud

Deliverable

production ML models | infrastructure

Required skills

GPU observability (DCGM, NVML), Systems programming (Go/Rust), Kubernetes controller development, LGTM stack (Loki, Grafana, Tempo, Mimir), SLURM cluster management, GitOps (ArgoCD), Multi-tenant isolation design, SRE practices.

Preferred skills

AI workload monitoring (NCCL, distributed training), Bare-metal systems administration, Cost analytics engineering.

Technologies

DCGM, NVML, Go, Rust, Kubernetes, Prometheus, Loki, Grafana, Tempo, Mimir, Thanos, VictoriaMetrics, SLURM, ArgoCD, OpenTelemetry, systemd.

Responsibilities

Build custom Prometheus exporters for GPU metrics; Develop Kubernetes controllers for GPU workload management; Deploy and manage LGTM stack for production-scale observability; Integrate SLURM clusters with Kubernetes for HPC workloads; Design multi-tenant observability isolation and RBAC; Implement intelligent alerting and SLO tracking for GPU failures and performance degradation.

Seniority

Senior, hands-on IC

Sourced via workable · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workable ↗