Site Reliability Engineer
Core
Partner with product and platform development teams to improve the stability, resilience, and operational readiness of cloud-based services for industrial asset management.
Role type
Site Reliability Engineer (SRE)
Builds
Cloud-native platform for reliability, observability, and developer autonomy in industrial environments
Domain
Industrial IoT, Cloud Infrastructure, Asset Management
Deliverable
production ML models | infrastructure
Required skills
Observability practices in distributed systems, SRE concepts (SLOs, error budgets, incident management), Cloud-native platforms, Infrastructure-as-code, Programming (TypeScript/Node.js preferred)
Preferred skills
Experience operating production systems, Mentoring developers on reliability practices
Technologies
TypeScript, Node.js, Cloud-native platforms, Infrastructure-as-code tools
Responsibilities
Assess service maturity and provide insights to development teams, Partner with development teams to implement observability best practices, Enable development teams to become autonomous with service deployment and support, Mentor developers on reliability practices, Act as a bridge to drive tooling and practice adoption across development teams, Contribute to company-wide initiatives defining reliability software development standards
Seniority
Mid-level, hands-on IC
