Staff Software Engineer, Data Products
Core
Design and build scalable, reusable data products and feature pipelines that power machine learning models for chronic disease management and personalized member experiences.
Role type
Staff Software Engineer, Data Engineering (ML Data Platform)
Builds
Production-grade feature datasets, batch and streaming pipelines, and self-service data foundations for ML teams.
Domain
Healthcare technology / Machine Learning Data Infrastructure
Deliverable
production ML models
Required skills
Large-scale distributed data systems design, Feature engineering, Cloud-native data platforms (AWS), Batch and streaming pipeline development, Data modeling (dimensional/event), CI/CD and observability, Technical leadership, Cross-functional stakeholder management
Preferred skills
Feature stores, MLOps workflows, Lakehouse architecture (Databricks/Iceberg), Streaming technologies (Kafka/Flink), NoSQL databases, Healthcare industry experience, Data/AI Governance
Technologies
Python, SQL, Apache Spark, AWS, Databricks, Iceberg, Kafka, Flink, Airflow, Redshift, Snowflake, Postgres, Docker, Kubernetes
Responsibilities
Design and maintain reusable feature datasets supporting personalization, risk prediction, and recommendation systems; Establish self-service foundations for dataset creation; Partner with Data Scientists to translate modeling requirements into production pipelines; Optimize large-scale distributed processing for performance and cost; Mentor engineers on distributed data processing and scalable modeling; Lead architecture discussions for ML data systems.
Seniority
Staff, technical leadership & cross-team influence