Site Reliability Engineer (SRE) I
Core
Support tools, processes, and operational practices to detect, investigate, respond to, and prevent production reliability issues using observability platforms and AI-enabled tools.
Role type
Junior Site Reliability Engineer (hands-on IC)
Builds
Reliable production services, operational tooling, dashboards, runbooks, and automation
Domain
Cloud infrastructure, distributed systems, observability, and production operations
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud infrastructure, distributed systems, observability and monitoring, scripting/programming, incident response, operational documentation, production telemetry, CI/CD, containers, infrastructure automation
Preferred skills
Kubernetes, infrastructure-as-code, Service Level Objectives, 24/7 on-call rotation, AI agent workflows, blameless post-incident reviews
Technologies
Python, Bash, PowerShell, JavaScript, Java, Go, Datadog, Dynatrace, New Relic, Splunk, Grafana, Prometheus, Elastic
Responsibilities
Support and maintain SRE operational tooling including dashboards, alerts, runbooks, and telemetry; investigate service-health issues using logs, metrics, and traces; participate in incident response and root-cause analysis; contribute to automation and integration work; review and validate operational artifacts; partner with engineering teams to identify reliability improvements.
Seniority
Junior, hands-on IC