Senior Service Reliability Engineer
Core
Ensuring stability and scalability of cloud gaming services and production infrastructure for PlayStation.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud gaming platform, large-scale web services, and production infrastructure for console-quality games.
Domain
Gaming, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Linux production systems engineering, Python, Bash, Go, Java, C++, or Rust, distributed data storage, NoSQL databases, data aggregation technologies, relational databases with high availability, monitoring and alerting, Kubernetes, AWS, software distribution, configuration management
Preferred skills
Software performance analysis, load testing, QA or SDET experience
Technologies
Hadoop, Ceph, MongoDB, Redis, Cassandra, ElasticSearch, Kafka, PostgreSQL, MySQL, Prometheus, Grafana, Kubernetes, AWS, Ansible, SaltStack, Puppet, Chef
Responsibilities
Lead ongoing improvements in reliability and scalability, define KPIs and processes for continuous improvement, influence architecture and implementation of solutions, mentor junior SRE staff, represent SRE in the wider organization, lead small-scale projects from inception to implementation, design platform-wide solutions
Seniority
Senior, hands-on IC with leadership responsibilities