Site Reliability Engineer
Core
Partner with product and platform development teams to improve the stability, resilience, and operational readiness of MaintainX's services.
Role type
Site Reliability Engineer (SRE)
Builds
Shared tooling for service deployment, support, and infrastructure; observability standards and incident response practices
Domain
Industrial operations platform (mobile-first work execution) / Cloud-native distributed systems
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
SRE concepts (SLOs, error budgets, incident management), cloud-native platforms, infrastructure-as-code, distributed system observability, mentoring developers
Preferred skills
TypeScript, Node.js
Technologies
Cloud-native platforms, Infrastructure-as-code tools
Responsibilities
Assess service maturity and provide insights to development teams; Partner with development teams to implement observability best practices; Enable development teams to become autonomous with their service deployment, support, and infrastructure; Mentor developers on reliability practices; Act as the bridge for Platform Division teams to drive tooling and practice adoption
Seniority
Mid-level, hands-on IC