Site Reliability Engineer
Core
Build and maintain the reliability, scalability, and resilience of Trainline's cloud-native platform serving millions of rail travelers across Europe.
Role type
Mid-level Site Reliability Engineer (SRE)
Builds
Observable, reliable, and scalable AWS-hosted infrastructure and shared platform services for a global rail booking platform.
Domain
Travel technology, Cloud Infrastructure, Site Reliability Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE concepts (SLI, SLO, error budgets), observability tooling (New Relic, ELK, Grafana), AWS cloud services, Linux troubleshooting, scripting (Python), infrastructure-as-code (Terraform), CI/CD (GitHub Actions), application architecture concepts (circuit breakers, throttling, health checks), time series data management.
Preferred skills
Experience with load balancing, reverse proxies, upstream configuration, worker & data flow concepts.
Technologies
AWS, New Relic, ELK stack, Grafana, Incident.io, Docker, ECS, Terraform, GitHub Actions
Responsibilities
Participate in production incident response and post-incident reviews; design and maintain observability using metrics, logs, and traces; improve monitoring and alerting to reduce noise and improve MTTD; support AWS-hosted infrastructure using IaC and CI/CD; collaborate with product engineering teams to ensure operational readiness; write reliable code and scripts for reliability goals.
Seniority
Mid-level, hands-on IC