Senior Site Reliability Engineer
Core
Define and maintain reliability targets, observability, and automation for a cloud-native vacation rental platform serving property owners globally.
Role type
Senior Site Reliability Engineer (IC)
Builds
Cloud infrastructure, Kubernetes clusters, and shared services for the Lodgify platform
Domain
SaaS / Vacation Rental / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Cloud Infrastructure, Observability (Metrics/Logs/Traces), Incident Response, Automation (Python), Stateful Systems (Databases/Caches/Queues), SLO/SLI Design, Disaster Recovery
Preferred skills
Internal Developer Platform design, Cost optimization, Toil reduction
Technologies
Datadog, Prometheus, Grafana, Kubernetes, Python
Responsibilities
Define SLIs/SLOs and reliability targets; Build actionable observability with metrics, logs, and traces; Automate operational tasks to reduce manual intervention; Participate in on-call and coordinate incident response; Improve reliability of stateful systems like databases and caches; Execute disaster recovery drills.
Seniority
Senior, hands-on IC