Senior Site Reliability Engineer, Wikimedia Enterprise
Core
Design, develop, and maintain reliable, scalable, and highly available infrastructure for Wikimedia Enterprise API services and data feeds.
Role type
Senior Site Reliability Engineer
Builds
Wikimedia Enterprise data ingestion services and API infrastructure for third-party reusers
Domain
Non-profit knowledge ecosystem / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Infrastructure as Code (Terraform, Ansible), Cloud platforms (AWS, Azure, GCP), CI/CD and GitOps (GitLab, ArgoCD), Observability (Prometheus, OpenTelemetry), SRE principles (SLOs, SLIs, error budgets), Incident response and on-call, Automation and scripting (Python, Go)
Preferred skills
Event streaming platforms (Kafka, Kinesis), Data lake architectures (Iceberg, Flink, Spark), Continuous profiling tools, Open source contributions
Technologies
Kubernetes, AWS, Azure, GCP, GitLab, ArgoCD, Terraform, Ansible, Prometheus, OpenTelemetry, Kafka, Kinesis, Iceberg, Flink, Spark
Responsibilities
Define and improve Service Level Objectives (SLOs) and error budgets; Build and enhance observability systems (metrics, logs, distributed tracing); Drive reliability engineering practices including capacity planning and chaos testing; Improve developer experience via self-service infrastructure; Design and optimize CI/CD and GitOps workflows; Implement secure-by-default infrastructure; Optimize infrastructure cost and efficiency using FinOps principles; Mentor peers in technical and operational strength
Seniority
Senior, hands-on IC