Senior CloudOps Engineer
Core
Own the reliability, performance, and observability of CloudZero's real-time ingestion path (Kafka) and shared critical infrastructure across AWS, Azure, and GCP.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Real-time ingestion systems, reliability tooling, and production Python services
Domain
Cloud infrastructure (AWS/Azure/GCP) and event-driven data processing
Deliverable
production ML models | product features | infrastructure
Required skills
Python, Infrastructure as Code (CloudFormation/SAM), SLO definition, asynchronous event-driven systems, production debugging, observability instrumentation
Preferred skills
Chaos engineering, load testing, internal developer portal experience, LLM-backed tooling
Technologies
Kafka, MSK, AWS, Azure, GCP, CloudFormation, SAM, Python, Sumo Logic, Datadog, Prometheus, Splunk
Responsibilities
Define and own SLOs for cross-team critical paths, build reliability tooling (load generators, fault injection), automate deployments and scaling, instrument systems for proactive failure detection, partner with product engineering on resilient architecture
Seniority
Senior, hands-on IC