Site Reliability Developer (python/java) / SRE
Core
Ensuring reliability and security of production cloud environments, leading incident response, and driving operational excellence for customer-facing systems.
Role type
Senior Site Reliability Developer (SRE)
Builds
Production cloud environments (AWS, Azure, hybrid) with high availability and security
Domain
Cloud Infrastructure & Security
Deliverable
production ML models | infrastructure
Required skills
Python, Java, Go, CloudFormation, Terraform, Kubernetes, Docker, Jenkins, Elasticsearch, Flink, Chaos Engineering, Service Level Objectives (SLOs)
Preferred skills
Serverless, Spark, New Relic, Lambda
Technologies
AWS, Azure, Kubernetes, Docker, Terraform, CloudFormation, Jenkins, Elasticsearch, Flink, Spark, New Relic, Lambda, GitHub, Artifactory, Jira
Responsibilities
Leading large-scale production incident response and postmortems; Defining operational and security policies and standards; Guiding development teams on Service Level Agreements (SLAs) and indicators; Developing automation to debug and fix complex production issues; Championing security and operational best practices across global teams.