CareerPlanSign in

Senior CloudOps Engineer

Boston💼 Full-time🗓 2026-07-31 → 2026-09-25

Core

Own the reliability, performance, and observability of CloudZero's real-time ingestion path (Kafka) and shared critical infrastructure across AWS, Azure, and GCP.

Role type

Senior Site Reliability Engineer (Infrastructure)

Builds

Real-time ingestion systems, reliability tooling, and production Python services

Domain

Cloud infrastructure (AWS/Azure/GCP) and event-driven data processing

Deliverable

production ML models | product features | infrastructure

Required skills

Python, Infrastructure as Code (CloudFormation/SAM), SLO definition, asynchronous event-driven systems, production debugging, observability instrumentation

Preferred skills

Chaos engineering, load testing, internal developer portal experience, LLM-backed tooling

Technologies

Kafka, MSK, AWS, Azure, GCP, CloudFormation, SAM, Python, Sumo Logic, Datadog, Prometheus, Splunk

Responsibilities

Define and own SLOs for cross-team critical paths, build reliability tooling (load generators, fault injection), automate deployments and scaling, instrument systems for proactive failure detection, partner with product engineering on resilient architecture

Seniority

Senior, hands-on IC

Sourced via ashby · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.