Site Reliability Engineer
Core
Ensures reliability, availability, performance, and scalability of software systems through monitoring, automation, and incident management.
Role type
Site Reliability Engineer (IC)
Builds
Stable, scalable software infrastructure and CI/CD pipelines
Domain
Enterprise Software / DevOps / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Site Reliability Engineering, Systems Engineering, DevOps, Observability, Logging, Performance Monitoring, Scripting, CI/CD, Incident Management, Root Cause Analysis, Troubleshooting
Preferred skills
Cloud Platforms (AWS, Azure, GCP), Container Orchestration (Docker, Kubernetes), Infrastructure as Code, Distributed Systems, Capacity Planning
Technologies
AWS, Azure, GCP, Docker, Kubernetes, CI/CD tools, Observability tools, Scripting languages
Responsibilities
Monitor system availability, performance, and reliability; Develop and maintain automation scripts and IaC solutions; Support CI/CD pipelines; Respond to and resolve system incidents; Troubleshoot technical issues; Collaborate with dev and ops teams to improve performance and minimize risk
Seniority
Mid-level, hands-on IC