Service Reliability Engineer - Applications Support
Core
Configure, tune, and fix multi-tiered systems to ensure optimal application performance, stability, and availability for data processing jobs.
Role type
Senior IC Site Reliability Engineer (Java/Big Data)
Builds
Self-healing systems, monitoring tools, and high-performance alerting for low-latency applications
Domain
Cloud infrastructure & Big Data processing
Deliverable
production ML models | infrastructure
Required skills
Java application debugging, Python programming, Kubernetes, AWS, Spark, Flink, Linux/Unix system administration, SRE principles, on-call management
Preferred skills
Hadoop, geographically distributed team collaboration, high-level project migrations
Technologies
Java, Python, Spark, Flink, Kubernetes, AWS, GCP, Hadoop
Responsibilities
Support Java-based applications and Spark/Flink jobs on bare-metal, AWS, and Kubernetes; Build automation for self-healing systems; Monitor production, staging, test, and development environments; Troubleshoot application-specific, core network, system, and performance issues.
Seniority
Senior, hands-on IC