SRE DevOps Engineer
Core
Ensuring scalability, availability, and reliability of applications and infrastructure through performance testing, monitoring, and chaos engineering.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production-grade scalable and resilient systems for enterprise clients
Domain
Cloud infrastructure, observability, and application performance management
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud monitoring (Splunk, SignalFx, ELK, AppDynamics, OpenTelemetry), AWS services (CloudTrail, CloudWatch, VPC Flow Logs, X-Ray), Python scripting, API development (REST, GraphQL), CI/CD tools (Jenkins, Git, Artifactory), containerization (Docker, Kubernetes), Infrastructure as Code (Terraform), chaos engineering, SLI/SLO definition
Preferred skills
Large-scale observability platform implementation, 24/7 production environment experience
Technologies
Splunk, SignalFx, ELK Stack, AppDynamics, ITRS, OpenTelemetry, AWS CloudTrail, CloudWatch, VPC Flow Logs, X-Ray, CloudWatch Insights, Python, REST, GraphQL, AWS SDK, Jenkins, Harness, Git, Artifactory, Docker, Kubernetes, Terraform
Responsibilities
Conduct performance testing and optimization for applications and infrastructure; Monitor system reliability and implement chaos engineering practices; Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets; Provide debugging and support for applications with a thorough understanding of REST APIs
Seniority
Senior, hands-on IC