Site Reliability Engineer (SRE)
Core
Implement, measure, and gather insights from Operational Level Indicators to identify service improvements covering availability, performance, resilience, incidents, and chronic problems while automating repetitive tasks.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Critical systems requiring high trust, resilience, and security
Domain
Financial services / Real-time streaming data / Cloud infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Apache Kafka event management, Splunk/Dynatrace monitoring, CI/CD pipeline management (Jenkins/Gradle/Maven/Bitbucket), Java code reading and debugging, RESTful services understanding, SQL query writing, microservices architecture in Cloud (GCP/Azure), UNIX shell scripting, Python scripting, disaster recovery exercises, technical vulnerability assessments
Preferred skills
Confluent Certified Administrator for Apache Kafka
Technologies
Apache Kafka, Splunk, Dynatrace, Jenkins, Gradle, Maven, Bitbucket, Java, SQL, GCP, Azure, Python, UNIX shell
Responsibilities
Implement jobs per Runbooks and automate repetitive tasks to reduce toil, lead and perform Disaster Recovery (DR) exercises, perform technical vulnerability assessments and recommend remediation, manage communication of production releases and service availability impact to stakeholders, troubleshoot major incidents and manage problem management, create SQL queries and write scripts using UNIX shell or Python
Seniority
Senior, hands-on IC