NOC Engineer / SRE
Core
24/7 service reliability, incident response, operational automation, and observability at the intersection of traditional NOC and engineering-driven reliability practices.
Role type
Senior Individual Contributor SRE/NOC Engineer
Builds
Self-healing systems, automated remediation scripts, and improved incident response workflows
Domain
Cloud infrastructure, network operations, and reliability engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux systems administration, incident management, cloud infrastructure (AWS), containers and orchestration (Docker, Kubernetes), scripting (Python, Bash, Go), networking fundamentals (DNS, TCP/IP, load balancing)
Preferred skills
SLO/SLI definition, Infrastructure as Code (Terraform, Ansible), security/compliance experience
Technologies
Grafana, Prometheus, Datadog, Splunk, CloudWatch, AWS, Azure, GCP, Kubernetes, Docker, Python, Bash, Go, Terraform, Ansible
Responsibilities
Act as primary or escalation responder in 24x7 on-call rotation, lead Major Incident response including triage and mitigation, design and maintain alerting strategies aligned with SLIs/SLOs, automate repetitive operational tasks to reduce toil, troubleshoot Linux-based systems and cloud platforms, support capacity planning and production release readiness
Seniority
Senior, hands-on IC