2024_MS_EDE3_XC_SRE_DataEngineering
Core
Design and engineer highly scalable, high-availability systems for data engineering projects, ensuring reliability, performance, and continuous deployment.
Role type
Senior Site Reliability Engineer (Data Engineering)
Builds
High-throughput data engineering services and infrastructure on Azure
Domain
Cloud Infrastructure / Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Azure cloud platforms, Kubernetes administration, Infrastructure as Code (Terraform, ARM, YAML), CI/CD frameworks, containerization (Docker, k8s), monitoring (ELK stack, Prometheus), security (DevSecOps), Python, Go
Preferred skills
Data pipelines, messaging systems, NoSQL databases, cloud cost optimization
Technologies
Azure, Terraform, ARM, YAML, Docker, Kubernetes, Argo, Flux, Helm, Istio, Grafana, Kustomize, ELK stack, Prometheus, Ansible
Responsibilities
Design scalable high-availability systems for data workloads; develop and manage monitoring systems with active alerting; automate deployments and policy enforcement; optimize system performance and resolve bottlenecks; manage infrastructure via code; implement security policies; plan scaling and cost management; handle system outages and perform root cause analysis.
Seniority
Senior, hands-on IC