Senior Staff Cloud Backend Engineer
Core
Design, implement, and maintain observability solutions and reliability components for large-scale datacenter infrastructure.
Role type
Senior Staff Cloud Backend Engineer (SRE/Observability)
Builds
Operational and reliability components of a large-scale Observability and Telemetry collection platform
Domain
Cloud Infrastructure / Datacenter Operations / Observability
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Backend software development, Cloud environments (AWS), Distributed systems, Observability tools (Prometheus, Grafana, ELK), SRE practices, Kubernetes, Docker, Terraform, Python, Go, Bash
Preferred skills
Multi-cloud platforms (AWS, Azure, GCP), Resource utilization optimization, Hardware/software integration
Technologies
AWS, Kubernetes, Docker, Terraform, Prometheus, Grafana, ELK Stack, Python, Go, Bash
Responsibilities
Implement SRE best practices for reliability and scalability; Develop automation scripts for infrastructure provisioning and management; Design and maintain observability solutions for datacenter infrastructure; Analyze and optimize performance of datacenter systems; Ensure compliance with security policies for observability solutions; Provide troubleshooting support for observability and reliability issues
Seniority
Senior Staff, hands-on IC