Site Reliability Engineer
Core
Ensuring the stability, scalability, and operational readiness of cloud gaming services and production infrastructure for the PlayStation brand.
Role type
Senior Site Reliability Engineer (Cloud Gaming)
Builds
Cloud gaming platforms and large-scale web services for console-quality video games
Domain
Gaming / Cloud Infrastructure / Distributed Systems
Deliverable
production ML models | product features | infrastructure
Required skills
Linux production systems administration, Python, Bash, Go, Java, C++, or Rust, distributed data storage at scale, NoSQL databases, data aggregation technologies, relational database management with high availability, monitoring and alerting, Kubernetes, AWS, software distribution, configuration management
Preferred skills
S/W performance analysis and load testing
Technologies
Hadoop, Ceph, MongoDB, Redis, Cassandra, ElasticSearch, Kafka, PostgreSQL, MySQL, Prometheus, Grafana, Kubernetes, AWS, ansible, saltstack, puppet, chef
Responsibilities
Lead team technical discussions on reliability and scalability, create high-level designs for new products and platforms, mentor junior SRE staff, lead incident response and post-mortem activities, collaborate on reliability improvements to address technical debt, implement automation to reduce toil
Seniority
Senior, hands-on IC