Site Reliability Engineer
Rewrite
## About the role
- Acts as a Designated Responsible Individual (DRI) working on call to monitor service for degradation, downtime, or interruptions.
- Alerts stakeholders as to the status and gains approval to restore system/product/service for simple problems.
- Contributes to efforts to collect, classify, and analyze data with little oversight on a range of metrics (e.g., health of the system, where bugs might be occurring).
- Contributes to the refinement of product features by escalating findings from analyses to inform decisions regarding the engineering of products.
- Contributes to the development of automation within production and deployment of a complex product feature.
- Contributes to efforts to ensure the correct processes are followed to achieve a high degree of security, privacy, safety, and accessibility.
- Implements solutions and mitigations to more complex issues impacting performance or functionality of Live Site service and escalates as necessary.
## Requirements
- Master's Degree in Computer Science, Information Technology, or related field AND 1+ year(s) technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.
- 1+ years experience managing physical infrastructure.
- Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience in software engineering, network engineering, or systems administration OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience in software engineering, network engineering, or systems administration OR equivalent experience.
- 2+ years technical experience working with large-scale cloud or distributed systems.
- 1+ year(s) people management experience.
- Experience working on large-scale distributed services with on-call responsibilities.
- Ability to build and influence broadly towards common goals and priorities.
- Ownership of end-to-end project lifecycle with solid project management and communication skills.
- Experience with managing physical infrastructure, supporting GPUs and InfiniBand.
## Nice to have
- (None specified in the original text)
## What we offer
- (None specified in the original text)
## About us
- (None specified in the original text)
Sourced via microsoft · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.