Site Reliability Engineer - Architect/Principal
Core
Design, implement, and maintain highly available, scalable, and secure systems for the IRIS Smart Manufacturing Platform across cloud and on-premise environments.
Role type
Principal Site Reliability Engineer (Architect)
Builds
Large-scale systems on AWS (EKS) or Azure (AKS) supporting hybrid deployments for mining, oil & gas, chemicals, and manufacturing.
Domain
Cloud infrastructure, Kubernetes, and Smart Manufacturing
Deliverable
production ML models | infrastructure
Required skills
Site Reliability Engineering, Linux systems administration, Kubernetes, Docker, CI/CD pipelines, Infrastructure as Code (Terraform, CloudFormation, AWS CDK), Service Mesh (Istio, Linkerd), observability (Prometheus, Grafana, EFK), security vulnerability management, root cause analysis, FinOps.
Preferred skills
Relational and NoSQL databases (Postgres, ElasticSearch, Redis), web server configuration (Nginx), messaging frameworks (Event Hub, Kafka, RabbitMQ).
Technologies
AWS, Azure, EKS, AKS, Kubernetes, Docker, Terraform, Ansible, Helm, Istio, Linkerd, Prometheus, Grafana, EFK, Nginx, Kafka, RabbitMQ, Postgres, ElasticSearch, Redis.
Responsibilities
Lead design and deployment of large-scale systems; serve as principal architect for reliability and disaster recovery; troubleshoot production issues and perform RCA; install and govern Service Mesh; implement SRE best practices (SLIs, SLOs, error budgets); optimize platform performance, scalability, and cloud cost; mentor junior engineers.
Seniority
Principal, hands-on IC with strategic mentorship