Senior Site Reliability Engineer
Core
Design, build, and maintain core infrastructure for developer focus, platform resilience, and scalability.
Role type
Senior Site Reliability Engineer
Builds
Production platforms and services for a global AI data annotation company
Domain
Cloud Infrastructure & DevOps
Deliverable
infrastructure
Required skills
AWS cloud architecture, Infrastructure as Code (Terraform/CloudFormation), CI/CD pipelines, monitoring and alerting, containerization (Docker), shell scripting
Preferred skills
MLOps operations, advanced web security principles, Go/Node.js/Python programming
Technologies
AWS (EC2, CloudFormation, ECS Fargate, Lambda, SQS, SNS, S3, ECR, RDS, Route 53), Terraform, CloudFormation, Serverless Framework, Bash, Python, Go, Grafana, ELK stack, CloudWatch, Prometheus, Docker, GitHub Actions
Responsibilities
Design and maintain core infrastructure; establish IaC best practices; troubleshoot and restore services; create run books for on-call resolution; deploy and monitor services; manage environment capacity; define availability/SLA standards; maintain monitoring and debugging tools; collaborate on performance improvements.
Seniority
Senior, hands-on IC