Analista de SRE III
Core
Ensuring reliability, performance, and security of critical production workloads in cloud environments.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Scalable, resilient, and secure distributed systems on AWS
Domain
Cloud Infrastructure & Data Technology
Deliverable
production ML models | infrastructure
Required skills
AWS, Kubernetes, Terraform, CI/CD pipelines, observability (metrics/logs/traces), incident management, Linux, networking, Python, Go, Bash, database management (PostgreSQL, MySQL, MongoDB, Redis, Elasticsearch, DynamoDB), security (IAM, secrets management), capacity planning, chaos engineering
Preferred skills
EKS, Helm, ArgoCD, Flux, OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, ELK/OpenSearch, SLI/SLO/SLA definition
Technologies
AWS, Kubernetes, Terraform, Ansible, Python, Go, Bash, PostgreSQL, MySQL, MongoDB, Redis, Elasticsearch, OpenSearch, DynamoDB, EKS, Helm, ArgoCD, Flux, OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, ELK, OpenSearch
Responsibilities
Architect, operate, and troubleshoot production workloads; manage CI/CD pipelines and GitOps practices; perform incident response, root cause analysis, and preventive action planning; monitor system health via observability tools; automate infrastructure and operations tasks; ensure system security and compliance; plan for capacity, performance, and scalability; conduct chaos engineering and resilience testing.
Seniority
Senior, hands-on IC