Staff Reliability Engineer
Core
Design and implement software and product reliability test regimens for cloud hardware to ensure highest quality delivery to customers.
Role type
Staff Reliability Engineer (Cloud Infrastructure Hardware)
Builds
Cloud hardware and data center servers meeting specified use-conditions and stresses
Domain
Cloud infrastructure, hardware reliability, data centers
Deliverable
production ML models | product features | dashboards & analysis | research | client delivery | infrastructure | physical/clinical work
Required skills
Design for Reliability principles, statistical & probability techniques, reliability modeling, Python, Unix (Bash/PowerShell), computer systems/hardware structure, switch/network interfaces
Preferred skills
Computer architecture, server architecture (block level), Hardware/Firmware/OS interactions, PCBA design/fabrication/validation, ReliaSoft & JMP statistical software, electronic components/devices failure modes, industry standards (IPC, JEDEC, Telcordia, MIL-STD)
Technologies
Python, Unix, Bash, PowerShell, ReliaSoft, JMP
Responsibilities
Create or revise reliability engineering guidelines to improve product field performance; use performance evaluation and prediction to improve reliability and maintainability of Cloud Infrastructure servers; identify, collect, analyze, and manage data to minimize failures; develop scripts representing expected environment and operational conditions; interface with program management, vendors, and design engineering on key reliability programs/issues.
Seniority
Staff, hands-on IC with strategic influence
