Site Reliability Engineer - India
Core
Build engineering discipline combining software and systems to develop creative solutions for operations problems, focusing on running better production applications and systems.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Automated operational work, self-healing and resiliency patterns, automated software and product upgrades, change and release management solutions
Domain
Enterprise mobile security, cloud infrastructure, container systems
Deliverable
production ML models | product features | infrastructure
Required skills
Software design, coding, and testing; Infrastructure components (routers, load balancers, cloud products, containers, compute, storage, networks); Debugging and troubleshooting; DevOps and application development; Large-scale software development (Java, Python, scripting); Kubernetes, Docker, Docker Swarm; Continuous Delivery tools; Unix (Linux, Solaris); Orchestration and configuration management; Data warehousing or big data environments
Preferred skills
Cross-domain expertise to solve complex mission-critical problems
Technologies
Java, Python, scripting languages, Kubernetes, Docker, Docker Swarm, DataDog, Unix/Linux/Solaris
Responsibilities
Design, code, test, and deliver software to automate manual operational work; Troubleshoot priority incidents and facilitate blameless post-mortems; Engage with development teams to develop software for reliability and scale; Identify application patterns and analytics for service level objectives; Design self-healing and resiliency patterns; Mentor and guide junior developers
Seniority
Senior, hands-on IC