CareerPlanGet AI match score →

Senior/Staff Site Reliability Engineer - Data Center

PathAI Boston🌐 Remote💼 Full-time💰 $165,750–$165,750🗓 2026-06-03 → 2026-07-31

Core

Designing, building, and operating a hybrid cloud/on-prem environment and data center to support a rapidly growing Machine Learning team.

Role type

Senior/Staff Site Reliability Engineer (Data Center & Hybrid Cloud)

Builds

Hybrid cloud infrastructure, data center facilities, and automation patterns for ML workloads

Domain

Healthcare AI / Machine Learning Infrastructure

Deliverable

infrastructure

Required skills

Infrastructure as Code (Terraform/CloudFormation), Python/Go scripting, Ansible, Datadog/Grafana/Prometheus, physical hardware administration (iDRAC/IPMI/Nvidia UFM/Juniper), high-performance storage optimization (Quobyte/S3/FSx/EFS), network operations across layers, container orchestration (EKS/ClusterAPI/KVM), incident response, root-cause analysis

Preferred skills

None explicitly stated

Technologies

AWS, Ansible, Python, Go, Datadog, Grafana, Prometheus, Terraform, CloudFormation, iDRAC, IPMI, Nvidia UFM, Juniper Systems, Quobyte, S3, FSx, EFS, EKS, ClusterAPI, KVM

Responsibilities

Implement SRE best practices focusing on monitoring and automation; Engineer infrastructure patterns for AWS cloud environments; Design and operate data center facilities; Integrate on-premises environments with cloud infrastructure; Improve reliability through root-cause analysis; Participate in on-call rotations and incident response

Seniority

Senior/Staff, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗