Engenheiros(as) de Software (SRE) | Observability, Cloud & AI
Core
Design, implement, and evolve observability solutions (metrics, logs, tracing) for cloud and Kubernetes environments, focusing on reliability, performance, and end-to-end visibility.
Role type
Senior Site Reliability Engineer (SRE) / Platform Engineer
Builds
Production cloud environments, Kubernetes clusters, and automated observability pipelines
Domain
Cloud Infrastructure, Observability, AI-driven Operations
Deliverable
production ML models | infrastructure
Required skills
SRE practices, Cloud platforms (Azure/AWS/GCP), Kubernetes, Distributed systems, Python/Go, CI/CD, Infrastructure as Code, Incident management, FinOps
Preferred skills
Observability tools (Dynatrace/Prometheus/Grafana), Multi-cloud environments, AI for operations, Platform Engineering experience, Cloud/Kubernetes certifications
Technologies
Dynatrace, Prometheus, Grafana, Kubernetes, Azure, AWS, GCP, Azure DevOps, Python, Go
Responsibilities
Design and evolve observability solutions including metrics, logs, and tracing; Develop automations to reduce manual toil; Operate and sustain cloud and Kubernetes environments; Lead critical incidents as Incident Commander; Participate in application and infrastructure architecture discussions; Identify cost optimization opportunities via FinOps; Apply AI for automation and productivity
Seniority
Senior, hands-on IC