Principal Reliability Engineer - EDS
Core
Senior technical authority defining and implementing reliability, resilience, and performance strategies for enterprise data platforms, cloud infrastructure, and data pipelines.
Role type
Principal Reliability Engineer (Strategy & Architecture)
Builds
Highly reliable, performant, and cost-efficient cloud-based data platforms and pipelines across AWS and GCP.
Domain
Enterprise Data Services, Cloud Infrastructure, Data Engineering
Deliverable
production ML models | infrastructure
Required skills
Strategic vision for reliability engineering, architectural direction for distributed systems, SLO/SLI framework design, AI-driven automation (AIOps), observability architecture, leadership of cross-organizational initiatives, Python scripting, Infrastructure-as-Code (Terraform/CloudFormation), CI/CD governance.
Preferred skills
Experience in regulated financial services/insurance environments, data governance and lineage systems, machine learning for operations, LLM/prompt engineering for operational tooling.
Technologies
AWS, GCP, Snowflake, EMR, Hadoop/Spark, Kubernetes, Python, Terraform, CloudFormation, Prometheus, Grafana, Datadog, Splunk, Dynatrace, OpenTelemetry, AWS Bedrock, SageMaker, Vertex AI.
Responsibilities
Define enterprise-wide reliability engineering strategy and roadmaps; architect and evolve cloud platforms for data services; lead AI-driven automation for anomaly detection and autonomous remediation; establish SLO/SLI frameworks; mentor Staff/Senior engineers; represent Reliability Engineering in executive architectural reviews.
Seniority
Principal, strategy & mentorship