Site Reliability Engineer II
Core
Own reliability of production domains for multi-tenant SaaS platforms (Encore, LabelTraxx) on AWS, designing observability/automation frameworks and acting as incident commander.
Role type
Senior individual contributor Site Reliability Engineer (SRE)
Builds
Monitoring, alerting, and incident response systems for AWS-hosted SaaS products
Domain
Enterprise software, financial systems, supply chain operations
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS (EC2, ECS/EKS, RDS, VPC), Terraform, CI/CD (GitHub Actions), Python/Bash scripting, OpenTelemetry/CloudWatch/Prometheus/Grafana, SLO/SLI management, incident response, networking, cloud security, AI-assisted engineering validation
Preferred skills
AWS Certified SysOps/DevOps, CKA/CKAD, Terraform Associate, multi-tenant SaaS architecture, PagerDuty, AI/LLM service operations
Technologies
AWS, Terraform, GitHub Actions, OpenTelemetry, CloudWatch, PagerDuty, ECS Fargate, EKS, Lambda, RDS PostgreSQL, OpenObserve
Responsibilities
Build and maintain monitoring/alerting tooling; analyze performance/capacity to maintain SLIs/SLOs; implement automated health checks and self-healing; lead incident response and post-mortems; build infrastructure automation in Terraform; develop CI/CD pipelines; operate containerized/serverless workloads; embed security (IAM, secrets) into automation; validate AI-generated code and tests
Seniority
Senior, hands-on IC