Site Reliability Engineer
Core
Architecting and deploying observability platforms to monitor system health, performance, and reliability while driving AI-driven alerting and proactive anomaly detection.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Autonomous operations, self-healing automation, and observability platforms
Domain
Cloud-native distributed systems & microservices
Deliverable
production ML models | infrastructure
Required skills
SRE principles, observability (Dynatrace, Datadog), Python scripting, Ansible, AWS, Azure, Docker, Kubernetes, CI/CD pipelines, chaos engineering, AI/ML for predictive analytics
Preferred skills
Large-scale SRE implementation, strategic mindset balancing engineering excellence with business priorities
Technologies
Datadog, Dynatrace, AWS, Azure, Python, Ansible, Docker, Kubernetes, Gremlin, Chaos Monkey
Responsibilities
Implement strategies for modernizing IT operations enhancing observability and toil reduction; Architect and deploy observability platforms to monitor system health, performance, and reliability effectively; Propose & drive strategies for AI-driven alerting and proactive anomaly detection to reduce MTTD & MTTR; Develop and enforce SRE best practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets; Establish & create AIOPS roadmap for improving operational efficiency; Lead efforts to automate repetitive tasks (toil) using scripting, orchestration tools, and AI/ML-based solutions; Drive incident management and root cause analysis processes through automation, ensuring continuous improvement to enable autonomous operations; Mentor and guide teams on adopting SRE principles and tools.
Seniority
Senior, hands-on IC with strategic mentorship