Senior Site Reliability Engineer - Workflow Automation
Core
Senior SRE owning reliability, scalability, and operational excellence of workflow orchestration platforms (Apache Airflow, Broadcom Automic/UC4), balancing steady-state operations with engineering projects to improve platform tooling and reduce toil.
Role type
Senior IC Site Reliability Engineer (Workflow Orchestration)
Builds
Workflow orchestration platforms (Airflow, Automic/UC4), automation tooling, observability solutions, and infrastructure-as-code components.
Domain
Data Engineering / Workflow Orchestration / Cloud Infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Apache Airflow (distributed executors, DAGs), enterprise job scheduling (Automic/UC4), Python, Linux/Windows systems, Kubernetes/Docker, CI/CD, observability (ELK, Grafana, Prometheus), Terraform/Helm/Ansible, incident management, SLO/SLI definition.
Preferred skills
AWS MWAA/Cloud Composer/Astronomer, dbt/Kafka/Snowflake, legacy scheduler migration, error budget management.
Technologies
Apache Airflow, Broadcom Automic/UC4, Kubernetes, Docker, Terraform, Helm, Ansible, ELK, Grafana, Prometheus, Python, AWS
Responsibilities
Serve as primary escalation point for production support involving Airflow and UC4; own and continuously improve SLOs, SLIs, and error budgets; monitor platform health and proactively remediate issues; partner with data engineering teams to troubleshoot DAG failures and scheduling issues; manage patching, upgrades, and configuration management; design and build automation tooling to reduce toil; develop and maintain infrastructure-as-code for platform components; build observability solutions for workflow visibility.
Seniority
Senior, hands-on IC