Especialista de SRE II
Core
Leading the SRE discipline for a critical company product, defining technical strategy, and managing the team to ensure high availability and operational efficiency.
Role type
Principal Site Reliability Engineer (Manager)
Builds
Critical company products and platforms
Domain
Technology / Data / Financial Services
Deliverable
production ML models | product features | infrastructure
Required skills
Kubernetes, Docker, AWS, Terraform, CI/CD tools, Observability platforms, Linux, Python/Go/Shell scripting, Infrastructure as Code, Performance Engineering, Disaster Recovery, Chaos Engineering
Preferred skills
AWS Professional/Specialty certification, Kubernetes CKA/CKS certification, Service Mesh (Istio), Backstage, Platform Engineering, FinOps, Event-driven architecture (Kafka, RabbitMQ)
Responsibilities
Lead and manage the SRE team technically and managerially; Define and evolve reliability, observability, and operational excellence strategy; Act as primary technical reference for critical incidents; Collaborate with Architecture, Development, Security, and Product teams; Conduct operational rituals, incident reviews, and capacity planning; Promote automation and continuous improvement culture; Ensure adoption of SRE, DevOps, and Platform Engineering best practices; Prioritize technical debt and reliability initiatives with business areas.
Seniority
Principal, hands-on IC with management responsibilities