Senior Technical Duty Officer, Cloud Ops
Core
Senior Incident Commander leading critical production incidents and building automation tools for cloud operations.
Role type
Senior IC Site Reliability Engineer (Incident Commander)
Builds
Incident response tooling, automation workflows, and observability processes for global cloud services.
Domain
Cloud Operations / SRE / Incident Management
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Incident Command, Python for automation, Linux/Unix troubleshooting, multi-tier distributed systems, cloud environments (GCP/AWS/Azure), container orchestration (Kubernetes), networking (DNS/TLS/HTTP), SRE fundamentals (SLOs/SLIs/error budgets), observability (metrics/logs/traces), change management, mentoring
Preferred skills
24x7 NOC/GTOC experience, Prometheus-compatible observability, distributed tracing, synthetic monitoring, ChatOps, dependency mapping
Technologies
Python, Linux, Kubernetes, GCP, AWS, Azure, PagerDuty, Jira, Prometheus, Grafana, SignalFx, Catchpoint
Responsibilities
Lead critical and blocker incidents from identification to recovery; design and build automation tools to reduce operational overhead; partner with SRE teams to improve service resiliency and observability; lead daily change reviews to minimize risk; mentor team members and improve runbooks; translate incident learnings into engineering improvements.
Seniority
Senior, hands-on IC