CareerPlanSign in

Senior Technical Duty Officer, Cloud Ops

Redwood City Office💼 Full-time🗓 2026-08-24 → 2026-09-25

Core

Senior Incident Commander leading critical production incidents and building automation tools for cloud operations.

Role type

Senior IC Site Reliability Engineer (Incident Commander)

Builds

Incident response tooling, automation workflows, and observability processes for global cloud services.

Domain

Cloud Operations / SRE / Incident Management

Deliverable

production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work

Required skills

Incident Command, Python for automation, Linux/Unix troubleshooting, multi-tier distributed systems, cloud environments (GCP/AWS/Azure), container orchestration (Kubernetes), networking (DNS/TLS/HTTP), SRE fundamentals (SLOs/SLIs/error budgets), observability (metrics/logs/traces), change management, mentoring

Preferred skills

24x7 NOC/GTOC experience, Prometheus-compatible observability, distributed tracing, synthetic monitoring, ChatOps, dependency mapping

Technologies

Python, Linux, Kubernetes, GCP, AWS, Azure, PagerDuty, Jira, Prometheus, Grafana, SignalFx, Catchpoint

Responsibilities

Lead critical and blocker incidents from identification to recovery; design and build automation tools to reduce operational overhead; partner with SRE teams to improve service resiliency and observability; lead daily change reviews to minimize risk; mentor team members and improve runbooks; translate incident learnings into engineering improvements.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.