Staff Site Reliability Engineer
Core
Architect core reliability platforms, lead incident response, and drive SRE best practices for a mission-critical AI market intelligence platform.
Role type
Staff Site Reliability Engineer (hands-on IC with mentorship)
Builds
Core reliability platforms, AIOps automation, self-service tooling, and observability infrastructure
Domain
AI-driven market intelligence / SaaS
Deliverable
production ML models | infrastructure
Required skills
SRE/DevOps leadership, production SaaS scaling, Python/Go, cloud platforms (AWS/GCP/Azure), Kubernetes, networking fundamentals, monitoring/alerting, incident command, full-stack troubleshooting
Preferred skills
AIOps strategy, advanced observability (OTEL, continuous profiling), "You Build It, You Run It" culture implementation
Technologies
Prometheus, Grafana, Datadog, ELK, OTEL, Kubernetes, AWS, GCP, Azure, Python, Go
Responsibilities
Architect reliability paved paths and self-service tooling; Lead AI-driven reliability and AIOps strategy; Champion SRE culture across engineering; Act as Incident Commander during critical events; Advance end-to-end observability; Mentor engineers across SRE and product teams
Seniority
Staff, hands-on IC with significant mentorship and architectural influence