Software Engineer, Observability
Core
Operate and maintain logging and metrics infrastructure to ensure operational health and platform performance for Lyft's distributed systems.
Role type
Senior Infrastructure Engineer (Observability)
Builds
Logging and metrics platforms, automation tooling, and monitoring systems for Lyft's infrastructure.
Domain
Cloud Infrastructure / Observability
Deliverable
production ML models | product features | infrastructure
Required skills
Go or Python, Kubernetes, Prometheus, Grafana, Loki, AWS, SLO definition, incident response, capacity planning, documentation
Preferred skills
Open-source tracing frameworks, multi-cluster management, automated task identification
Technologies
Go, Python, Kubernetes, AWS, Prometheus, Grafana, Loki
Responsibilities
Maintain and improve tooling for reliability and scalability; define SLOs and provide monitoring tooling; analyze metrics for fault detection; collaborate on observability alignment and capacity planning; maintain infrastructure documentation; automate repetitive tasks; participate in on-call rotations and incident response.
Seniority
Senior, hands-on IC