Senior Data Engineer
Core
Design, implement, and support automation solutions and data pipelines for food processing, powering real-time dashboards, ML training datasets, and operational analytics.
Role type
Senior IC data engineer (real-time & batch)
Builds
Scalable ETL/ELT pipelines, real-time ingestion workflows, data models, and automation for labeling/CV pre-annotation
Domain
Food processing industry + distributed data systems
Deliverable
production ML models | product features | dashboards & analysis
Required skills
Python, SQL, distributed compute frameworks (Spark/Ray/Airflow), streaming systems (Kafka/NATS/Redis), AWS services, NoSQL/Columnar DBs (PostgreSQL/MongoDB/ClickHouse), data modeling, data governance
Preferred skills
NATS JetStream, GPU-based CV pipelines, ClickHouse Materialized Views, HashiCorp ecosystem (Nomad/Consul/Vault)
Technologies
Python, SQL, Apache Spark, Ray, Apache Airflow, Kafka, NATS JetStream, Redis Streams, AWS (S3, Lambda, EC2, Glue), ClickHouse, PostgreSQL, MongoDB
Responsibilities
Architect scalable ETL/ELT pipelines for batch and streaming; Design real-time ingestion workflows with NATS JetStream; Develop data models for ClickHouse; Manage storage across AWS S3 and on-prem; Build automation for data labeling and versioning; Ensure data quality and lineage; Collaborate on ML training datasets; Define data governance policies; Monitor multi-region pipeline performance
Seniority
Senior, hands-on IC