Site Reliability Engineer
Core
Design, maintain, and optimize scalable cloud infrastructure to support the Tekmetric auto repair shop management platform.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Cloud infrastructure, automated pipelines, monitoring systems, and disaster recovery solutions for the Tekmetric SaaS platform.
Domain
SaaS / Auto repair industry / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Cloud infrastructure architecture, Infrastructure as Code (Terraform), Container orchestration (Kubernetes), CI/CD pipeline design, Observability (Prometheus, Grafana, ELK), Scripting (Python, Bash), Incident response, Cross-functional collaboration.
Preferred skills
Go, Java, Javascript, Compliance and security best practices.
Technologies
AWS, GCP, Terraform, Docker, Kubernetes, Prometheus, Grafana, ELK stack, Python, Bash, Go, Java, Javascript.
Responsibilities
Architect and maintain reliable, scalable, and secure cloud infrastructure; Develop and maintain monitoring, alerting, and incident response practices; Create automated pipelines for deployment, testing, and infrastructure management; Implement and manage solutions for backup, disaster recovery, and failover processes; Apply best practices in security, monitoring, and compliance; Work cross-functionally with development, data, product, and QA teams; Provide technical leadership and mentorship to junior DevOps team members.
Seniority
Senior, hands-on IC with mentorship responsibilities