Senior Reliability Developer
Core
Owns and maintains public/private cloud infrastructure, monitoring systems, and CI/CD pipelines to ensure resilience and security for business services.
Role type
Senior Reliability Developer (Infrastructure & SRE)
Builds
Resilient multi-region cloud infrastructure, automated monitoring solutions, and secure deployment pipelines for the Aurora Platform.
Domain
Cybersecurity / Cloud Infrastructure / DevOps
Deliverable
production ML models | infrastructure
Required skills
Infrastructure-as-Code (Terraform), Container Orchestration (Kubernetes, ECS, EKS), Cloud Architecture (AWS), Scripting (Python, Bash, JavaScript), Monitoring (Prometheus, Grafana, Zabbix), CI/CD (Jenkins, GitLab), Certificate Management, Distributed Systems Troubleshooting
Preferred skills
None stated
Technologies
Terraform, Puppet, Chef, Docker, Kubernetes, ECS, EKS, AWS (EC2, ECS, EKS, ELB, S3, RDS, IAM, Lambda, CloudFormation, VPC, Route53, CloudWatch), Prometheus, Grafana, Zabbix, AlertManager, PagerDuty, Jenkins, GitLab, GitHub, Bitbucket, Gaia, Keeper, Jira, Confluence, Backstage, LucidChart
Responsibilities
Design and manage multi-region cloud infrastructure deployments; Implement automation for infrastructure provisioning and patching; Create and improve service monitoring solutions; Troubleshoot microservices and container networking issues; Manage Kubernetes cluster operations and upgrades; Maintain CI/CD pipelines; Create and maintain technical documentation; Execute service decommissioning; Participate in on-call rotation for incident response and root cause analysis.
Seniority
Senior, hands-on IC