Senior Site Reliability Engineer
Core
Provide stability and operational excellence for Grab's core transport booking systems and high-throughput distributed services.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud infrastructure, automated solutions, and high-throughput streaming services for the Mobility team.
Domain
Ride-hailing / Superapp / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Go, Python, C, C++, Java, Perl, Ruby, Linux administration, shell scripting, Infrastructure as Code, service monitoring, log management, automation tools (Jenkins, Ansible, Chef, SaltStack, Puppet), system troubleshooting, algorithm design, data structures, complexity analysis.
Preferred skills
Golang, cloud infrastructure (AWS, Azure, GCP), containerization (Docker), container orchestration (Kubernetes), streaming processing (Flink), open source contribution, performance analysis.
Technologies
Go, Python, C, C++, Java, Perl, Ruby, Linux, Jenkins, Ansible, Chef, SaltStack, Puppet, AWS, Azure, Google Cloud Platform, Docker, Kubernetes, Flink.
Responsibilities
Provide automated solutions for manual tasks, manage cloud infrastructure using IaC, diagnose incidents and engage teams on resolutions, support long-term infrastructure decisions, drive operational excellence practices, lead and mentor junior engineers.
Seniority
Senior, hands-on IC