Data Engineer
Core
Design and optimize robust data pipelines and observability services for an internal analytics platform supporting AI data solutions and ML pipelines.
Role type
Senior Data Engineer (Analytics & Observability)
Builds
Scalable batch and real-time data pipelines, integrations with annotation tools and ML validation systems, observability alerts, and dashboards.
Domain
AI data solutions, machine learning infrastructure, cloud-native data platforms
Deliverable
production ML models | infrastructure
Required skills
Python, SQL, PySpark, AWS Glue, Airflow, data modeling, schema design, data governance, API design, Kafka, database fundamentals
Preferred skills
Apache Hadoop, Apache Spark, Delta Lake, Redshift, Snowflake, Athena, QuickSight, Redash, infrastructure-as-code, CI/CD
Technologies
AWS (S3, Lambda, Glue, Kinesis, Firehose, RDS), PySpark, Airflow, Kafka, Athena, QuickSight, Redash, Delta Lake, Redshift, Snowflake
Responsibilities
Design and build scalable batch and real-time data pipelines; Integrate analytics services with upstream annotation tools and downstream ML systems; Develop ETL/ELT workflows ensuring data quality and lineage; Implement observability pipelines and alerts for mission-critical metrics; Build data models and queries to power dashboards; Contribute to infrastructure-as-code and CI/CD practices; Document architecture and support runbooks.
Seniority
Senior, hands-on IC