Senior SRE
Core
Own operational excellence and reliability of modernized SaaS applications delivered to Operating Companies (OpCos) via the Banyan AI Factory.
Role type
Senior hands-on Site Reliability Engineer (SRE)
Builds
Production reliability and automation for distributed containerized applications on AWS and Azure
Domain
SaaS, Cloud Native, DevSecOps, AI-assisted Engineering
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Python/JavaScript/Go, Kubernetes, Terraform, AWS/Azure, CI/CD, Observability, Incident Response, Disaster Recovery, AI-assisted coding
Preferred skills
Wiz, Prisma Cloud, Checkov
Technologies
AWS, Azure, Docker, Kubernetes, Terraform, GitHub Actions, GitLab CI, Datadog, New Relic, Dynatrace, Claude Code
Responsibilities
Provide 24x7 on-call coverage for OpCo applications; manage cloud integrations and production health; implement observability tooling; execute disaster recovery procedures; respond to security incidents; automate environments using IaC and CI/CD; build and operate AI agents for SRE tasks; resolve complex infrastructure and automation issues
Seniority
Senior, hands-on IC