Senior Site Reliability Engineer, AIOPs
Core
Operate an AI Data Center AIOps platform that transforms high-volume telemetry into reliable insights and automation for GPU fleets.
Role type
Senior Site Reliability Engineer (AIOps Platform)
Builds
Telemetry ingestion, processing, storage, APIs, and dashboards for GPU fleet management
Domain
AI Data Centers, GPU Infrastructure, Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes operations, distributed systems, incident response, SLO/SLI management, infrastructure as code, Python scripting, CI/CD pipelines, observability stack expertise
Preferred skills
Linux networking fundamentals, streaming systems operations (Kafka/Pulsar, Flink/Spark), automation tool development, large-scale cluster management
Technologies
Kubernetes, Terraform, Helm, Python, Bash, Prometheus, Grafana, Kafka, Pulsar, Flink, Spark, ClickHouse, Elastic, TSDBs
Responsibilities
Monitor platform health via dashboards/logs/metrics and automate recurring checks, own Kubernetes deployments end-to-end including runbooks and rollbacks, lead first-level incident triage and root cause analysis, build and maintain runbooks/SOPs/checklists, manage deployment infrastructure and packaging
Seniority
Senior, hands-on IC