Site Reliability Engineer - Big Data (7 to 11 years)
Core
Managing and maintaining complex, distributed big data ecosystems to ensure reliability, scalability, and security of large-scale production infrastructure.
Role type
Senior Site Reliability Engineer (Big Data)
Builds
Reliable, scalable, and secure big data infrastructure and automation systems for PhonePe's digital payments and financial services.
Domain
Fintech / Big Data Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Linux administration, scripting (Perl, Golang, Python), Hadoop stack (HDFS, HBase, Airflow, YARN, Ranger, Kafka, Pinot), configuration management (Puppet, Salt, Chef, Ansible), networking, observability (ELK, Grafana, Prometheus, Open Telemetry), automation, incident response, capacity planning.
Preferred skills
Public cloud platforms (AWS, Azure, GCP), system architecture design, observability tool implementation.
Responsibilities
Manage and support Linux/Unix environments; lead on-call rotations and incident responses with root cause analysis; design and implement automation for big data infrastructure; troubleshoot complex production issues; design and review scalable system architectures; enforce security standards; develop tools to automate operational processes; monitor and optimize system performance and resource usage; collaborate with development teams on reliability best practices.
Seniority
Senior, hands-on IC