Sr. Staff Observability Software Engineer
Core
Design, implement, and maintain observability solutions (monitoring, alerting, logging, tracing) for hyper-scale distributed systems to ensure service health and reliability.
Role type
Sr. Staff Observability Software Engineer (IC)
Builds
Next-generation Observability Platform based on Kubernetes and OSS solutions
Domain
Cloud Infrastructure / Distributed Systems / Observability
Deliverable
production ML models | product features | infrastructure
Required skills
Large-scale distributed systems engineering, observability solution implementation, monitoring/alerting/logging systems, KPI/SLO definition, complex incident troubleshooting, observability standardization, technical mentorship
Preferred skills
Containerization and orchestration (Docker, Kubernetes), APM tools (Dynatrace, AppDynamics), distributed tracing (Jaeger, Zipkin), cloud infrastructure (AWS, Azure, GCP), DevOps/SRE practices, IaC, Go/Java/Python/Ruby proficiency
Technologies
Kubernetes, OpenTelemetry, Prometheus, Grafana, Elastic Stack, Docker, AWS, Azure, Google Cloud Platform, Jaeger, Zipkin, Dynatrace, AppDynamics
Responsibilities
Design and maintain observability solutions across platforms; collaborate with cross-functional teams to define requirements; develop best practices for telemetry systems; evaluate and recommend industry-leading tools; define and track KPIs and SLOs; assist in troubleshooting complex incidents; provide guidance and mentorship to engineers; conduct system evaluations and optimizations; drive standardization of observability processes; contribute to documentation and runbooks.
Seniority
Sr. Staff, hands-on IC with strategic influence