Senior Site Reliability Engineer, Infrastructure Foundations
Core
Senior SRE responsible for ensuring the reliability and delivery of Wikimedia's global top-10 website (Wikipedia) and its underlying open-source infrastructure.
Role type
Senior Site Reliability Engineer (Infrastructure)
Builds
Wikimedia's public-facing infrastructure and scalable services for millions of users globally
Domain
Non-profit / Open Source / Web Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Linux system administration, configuration management (Puppet), scripting (Python, Go, Bash), incident response, root cause analysis, automation, security incident response, SLO definition
Preferred skills
fleet-wide security policies, software supply chain security, credential management, immutable logging, monitoring (Prometheus, Grafana), LAMP stack, MediaWiki
Technologies
Puppet, Kubernetes, Python, Go, Bash, Ruby, Debian, Prometheus, Grafana, PHP, HHVM, memcached, Redis
Responsibilities
Perform day-to-day operational/DevOps tasks (deployment, maintenance, configuration, troubleshooting), implement configuration management and deployment tools, lead continuous improvement by automating service installation and maintenance, assist in architectural design of new services, participate in 24/7 on-call rotation for incident response, mentor peers
Seniority
Senior, hands-on IC with mentorship responsibilities