Consulting/Principal Site Reliability Engineer
Core
Design, build, and operate highly reliable and scalable infrastructure and observability systems across cloud environments to support the Brightmine Data Platform.
Role type
Principal Site Reliability Engineer
Builds
Cloud infrastructure, observability systems, CI/CD pipelines, and automation tools for the Brightmine Data Platform
Domain
Cloud infrastructure, data platform engineering, HR technology
Deliverable
production ML models | infrastructure
Required skills
Python, PowerShell, Shell, AWS, Azure, Terraform, Docker, Kubernetes, Git, Prometheus, Grafana, Linux/Unix administration
Preferred skills
Performance tuning, security best practices, migration strategies
Technologies
Terraform, GitHub, Docker, Kubernetes, Prometheus, Grafana, AWS, Azure
Responsibilities
Design and implement monitoring, alerting, and observability systems; Lead incident response activities; Design and maintain scalable cloud infrastructure solutions; Implement Infrastructure as Code (IaC); Build and maintain CI/CD pipelines; Optimize application and infrastructure performance; Develop automation tools and scripts; Implement security best practices
Seniority
Principal, hands-on IC