Staff Software Engineer I - SRE
Core
Expert-level SRE driving proactive reliability improvements and incident management for a multi-cloud data streaming platform.
Role type
Staff Software Engineer I - SRE
Builds
Automation, tooling, and reliability standards for Confluent Cloud across AWS, GCP, and Azure.
Domain
Cloud infrastructure, distributed systems, data streaming (Kafka), and incident management.
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE/incident management, multi-cloud (AWS/GCP/Azure), distributed systems, observability (metrics/logging/tracing), Kubernetes, CI/CD, SLO/SLA frameworks, Rootly/PagerDuty, systems thinking.
Preferred skills
Kafka/event streaming expertise, AI-assisted workflows, async collaboration across time zones.
Technologies
AWS, GCP, Azure, Kubernetes, Rootly, PagerDuty, Jira, Confluence, Slack, Kafka.
Responsibilities
Analyze systemic failure patterns and design reliability improvements; define and maintain SLO/SLA frameworks; build automation to reduce incident response toil; serve as on-call Incident Commander; coach teams through post-mortems; edit customer-facing incident documents.
Seniority
Staff, strategic program ownership with hands-on engineering