Site Reliability Engineer (SRE) – AWS + Docker
Core
Improve reliability, scalability, and performance of cloud-native systems via AWS infrastructure management, containerized workload operations, CI/CD enablement, and incident response.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Scalable, highly available AWS environments and containerized services
Domain
Cloud Infrastructure (AWS) & Container Orchestration
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
AWS (EC2, ECS/EKS, IAM, VPC, ALB/NLB, Route 53, S3, CloudWatch), Docker, Kubernetes/EKS or ECS, CI/CD (GitHub Actions, Jenkins, Azure DevOps), Infrastructure as Code (Terraform, CloudFormation), Python, Bash, Linux system administration, Networking (DNS, TCP/IP, TLS, security groups, NACLs), Observability (Prometheus/Grafana, ELK/OpenSearch, X-Ray)
Preferred skills
CloudFront, RDS, ElastiCache, Auto Scaling Groups, Blue/green and canary deployment strategies, Artifact management, Vulnerability scanning, Secrets management, Circuit breakers, retries, backoff, health checks
Technologies
AWS, Docker, Kubernetes, Terraform, CloudFormation, GitHub Actions, Jenkins, Azure DevOps, Prometheus, Grafana, ELK, OpenSearch, X-Ray, Python, Bash
Responsibilities
Define and maintain SLOs, SLIs, SLAs, and error budgets; Build and manage AWS infrastructure for scalable, highly available systems; Operate containerized services using Docker and ECS/EKS/Kubernetes; Implement and optimize CI/CD pipelines and deployment strategies; Establish observability through metrics, logs, and traces; Automate infrastructure and operations using IaC and scripting; Manage incident response, runbooks, root-cause analysis, and remediation; Drive performance tuning, capacity planning, and cost optimization; Implement security best practices across infrastructure and deployments; Partner with development teams to improve reliability by design
Seniority
Senior, hands-on IC