Service Reliability Engineer
Core
Ensure stability and operational readiness of large-scale cloud gaming services and production infrastructure.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud gaming platforms and production services for PlayStation
Domain
Gaming / Cloud Infrastructure
Required skills
Linux systems administration, Python, Bash, Go, Java, C++, Rust, Kubernetes, AWS, distributed data storage, NoSQL, data aggregation, RDBMS, monitoring & alerting, incident management, automation, load testing
Preferred skills
SDET experience, mentoring junior staff
Technologies
Hadoop, Ceph, MongoDB, Redis, Cassandra, ElasticSearch, Kafka, PostgreSQL, MySQL, Prometheus, Grafana, Ansible, SaltStack, Puppet, Chef
Responsibilities
Lead technical discussions on reliability and scalability, create high-level designs for new products, mentor junior SREs, lead incident response and post-mortems, collaborate on reliability improvements, contribute to code, implement automation to reduce toil