Network Operations Center (NOC) Analyst
Core
First-line technical response for 24/7 operations across high-performance compute data centers, focusing on telemetry analysis, infrastructure monitoring, and independent diagnosis of compute, network, and hardware systems.
Role type
NOC Analyst (Infrastructure Operations)
Builds
Stability and observability of large-scale AI compute infrastructure
Domain
Data Center Operations / High-Performance Computing (HPC)
Deliverable
Infrastructure monitoring and incident triage
Required skills
Linux command-line proficiency, TCP/IP networking concepts, system log analysis, independent technical diagnosis, incident triage and escalation, technical documentation
Preferred skills
Grafana, Datadog, Prometheus, HPC or GPU-based infrastructure experience, Bash or Python scripting
Technologies
Linux, TCP/IP, DNS, ping, traceroute, netstat, tcpdump, Grafana, Datadog, Prometheus
Responsibilities
Monitor data center systems using telemetry data and alerting tools to detect anomalies; Perform independent technical diagnosis across Linux systems, network connectivity, and hardware health; Troubleshoot network-layer issues including connectivity, routing, and interface errors; Triage and escalate incidents to appropriate teams with technically accurate summaries; Create and maintain detailed tickets documenting diagnostic steps and findings; Identify recurring alert patterns to improve monitoring coverage.
Seniority
Individual Contributor (IC)