Site Reliability Engineer
Core
Bridge software engineering and systems architecture to build resilient, automated cloud infrastructure and observability platforms for IT services.
Role type
Senior Site Reliability Engineer (Infrastructure & Observability)
Builds
Production-grade cloud infrastructure, CI/CD pipelines, observability stacks, and internal AI automation tools.
Domain
Cloud Infrastructure & DevOps
Deliverable
production ML models | infrastructure
Required skills
Python, Terraform, Kubernetes, AWS/Azure/GCP, CI/CD (GitHub Actions), Observability (Datadog/Prometheus/ELK), Distributed Systems (Kafka)
Preferred skills
Pulumi, Self-hosted runners, Incident management workflows
Technologies
Terraform, Pulumi, AWS, Azure, GCP, Kubernetes, Docker, GitHub Actions, Datadog, Prometheus, ELK, Kafka
Responsibilities
Design and deploy IaC infrastructure; Optimize system performance and scaling; Architect CI/CD pipelines; Enable logging, metrics, and alerts by default; Build internal AI plugins and automation scripts; Lead incident response and post-mortems; Collaborate with Security and Engineering teams.
Seniority
Senior, hands-on IC