Analista de SRE Pleno - Vaga Afirmativa para Mulheres
Core
Design, operate, and continuously improve the reliability, scalability, observability, and performance of cloud-native platforms supporting critical business applications, data pipelines, and AI/ML workloads.
Role type
Mid-level Site Reliability Engineer (SRE)
Builds
Cloud-native platforms, Kubernetes infrastructure, data processing environments, and AI/ML support systems
Domain
Cloud Infrastructure, Data Engineering, AI/ML Platforms
Deliverable
production ML models | infrastructure
Required skills
AWS services, Kubernetes, distributed systems troubleshooting, Datadog observability, Airflow, Amazon EMR, S3, Terraform, CI/CD pipelines, Linux, Python/Bash scripting
Preferred skills
Large-scale cloud-native operations, Kubernetes cluster lifecycle management, SRE principles (SLOs/SLIs/error budgets), Prometheus/Grafana/OpenTelemetry, Spark, AWS/Kubernetes/Terraform/Datadog certifications, technical English
Technologies
AWS, Kubernetes, Datadog, Airflow, Amazon EMR, S3, Terraform, Python, Bash, Linux, Prometheus, Grafana, OpenTelemetry, Spark
Responsibilities
Design and operate highly available, scalable, and secure cloud platforms on AWS; Build and maintain Kubernetes-based infrastructure; Improve platform reliability through automation and IaC; Implement and enhance Datadog observability solutions; Support and optimize large-scale data processing environments; Partner with Data and AI teams to improve ML platform maturity; Lead incident response, root cause analysis, and post-incident reviews; Define and measure SLIs, SLOs, and error budgets; Improve CI/CD deployment processes; Optimize cloud infrastructure utilization, performance, and costs; Mentor team members and promote SRE best practices
