Senior Site Reliability Engineer
Core
Build and operate a resilient, observable, and scalable data & ML platform enabling mission-critical workloads across the organization.
Role type
Senior Site Reliability Engineer (Data & ML Platform)
Builds
Production-grade data and ML infrastructure on hybrid cloud environments
Domain
Healthcare technology, Cloud Infrastructure, Data Engineering
Deliverable
infrastructure
Required skills
SRE mindset, Cloud infrastructure architecture, Databricks, Snowflake, Observability (Datadog), CI/CD automation, Infrastructure as Code (Terraform), Event-driven architecture, Python, Shell scripting
Preferred skills
DevSecOps, ML infrastructure tooling (MLflow, Feature Stores), Large-scale lakehouse architecture (Iceberg, Glue), Compliance-aware architecture (HIPAA, SOC 2), Multi-cloud experience (Azure), Open-source contributions
Technologies
Databricks, Snowflake, AWS, Datadog, GitHub Actions, Terraform, EventBridge, SNS/SQS, Lambda, Kafka, S3, Delta Lake
Responsibilities
Own Databricks & Snowflake platform lifecycle including automation and cost optimization; Architect resilient, scalable, and secure infrastructure; Build and maintain platform-wide monitoring, alerting, and logging; Automate deployments of data pipelines and ML workflows; Build patterns for inter- and intra-cloud data movement; Champion event-driven architectures; Serve as platform partner for analytics and data science teams; Influence engineering-wide decisions on data platform architecture.
Seniority
Senior, hands-on IC