Principal Site Reliability Engineer, Infrastructure Observability
Core
Formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the observability, sustainability, scalability, measurability, and recoverability of cloud & on-prem solutions.
Role type
Principal Site Reliability Engineer (Infrastructure Observability)
Builds
Cloud & on-prem solutions for an asset management organization
Domain
Asset Management / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Cloud infrastructure design and operations, DevOps practices, CI/CD toolchain, incident response, 24x7 monitoring, automation, chaos modeling, SRE methodologies, Service Level Objectives (SLOs) definition, Error Budget management, observability tooling, cloud management tools, multi-language programming, database development
Preferred skills
Cloud or SRE-related certifications, Azure working knowledge
Technologies
Amazon AWS, New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk, Ansible, Terraform, Vault, Vagrant, Python, Java, GO, Node.js, .Net Core, SQL Server, PostgreSQL, MySQL
Responsibilities
Design technology solutions to prevent or minimize service disruptions, foster a culture of deep learning through blameless post-mortems, transform operations teams by adopting SRE standard methodologies, analyze incidents impacting technology availability, drive initiatives to reduce or prevent technology failures, pull together information from disconnected systems into cohesive views, contribute to definition of target state architecture, develop or mentor diverse talent on the team
Seniority
Principal, hands-on IC with strategic and program-level implementation experience