Senior Site Reliability Engineer
Core
Design and implement automation solutions to improve service reliability, availability, and performance for enterprise customers while providing operational insights to product teams.
Role type
Senior Site Reliability Engineer (Customer Support & Reliability)
Builds
Automation tooling, proactive alerting systems, and service telemetry enhancements for large-scale distributed systems.
Domain
Cloud Infrastructure & Enterprise Support
Deliverable
production ML models | infrastructure
Required skills
Large-scale cloud services operations, Service Reliability/Availability/Performance improvement, Observability and MELT implementation, Logic Apps, Jupyter Notebooks, incident root cause analysis and automation, distributed systems troubleshooting, service telemetry design.
Preferred skills
None stated.
Technologies
Logic Apps, Jupyter Notebooks
Responsibilities
Collaborate with engineering teams to build automation for faster issue resolution; interface with enterprise customers for service escalations; design telemetry changes for automation consumption; analyze data to provide operational insights to product teams; influence product architecture for supportability.
Seniority
Senior, hands-on IC