CareerPlanGet AI match score →

Data Engineer

💼 Full-time🗓 2026-07-30

Core

Designing and maintaining scalable ETL pipelines and distributed data systems for machine learning and data science teams.

Role type

Data Engineer

Builds

ETL pipelines and data engineering infrastructure

Domain

Data Engineering / Machine Learning Infrastructure

Deliverable

production ML models

Required skills

Large-scale data processing (Spark, Hadoop, Airflow), Python/Scala/Go, Spark cluster management, Terraform, Docker, Kubernetes, Open-source contribution

Preferred skills

Experience with vehicle data, Convex optimization, Time-series analytics

Technologies

Spark, Hadoop, Airflow, Python, Scala, Go, Terraform, Docker, Kubernetes

Responsibilities

Design scalable ETL systems, identify and resolve scalability bottlenecks, manage Spark clusters, contribute to engineering infrastructure, improve data quality and discoverability

Seniority

Mid-level, hands-on IC

Rewrite
## About the role You are a thoughtful engineer. You understand the complexities of distributed systems and how to triage and solve issues that arise with them. Scalability is top of mind when designing any system or writing code. You believe building a better ETL system requires close collaboration with the machine learning and data science teams. You avoid reinventing the wheel unless necessary and are excited by opportunities to contribute to the open-source community. ## Onboarding Timeline ### Day 5 - Learn about Viaduct's history and mission - Get to know every team member - Set up your development environment - Understand Viaduct's ETL pipelines and run your first DAGs - Deep dive into the nuances of vehicle data - Attend our weekly ML lunch ### Day 30 - Take ownership of ETL pipelines - Identify scalability bottlenecks in the existing ETL pipelines - Be familiar with the day-to-day work of machine learning engineers and data scientists - Learn the architecture of data engineering systems and services ### Day 90 - Be the ETL pipeline expert at Viaduct - Improve overall data quality and discoverability - Confident in the scalability of Viaduct's ETL pipelines - Present your work at our weekly ML lunch - Comfortable contributing to our engineering infrastructure and systems ## Expected Skills - 2+ years working with large-scale data processing tools (Spark, Hadoop, Airflow, etc) - Expertise in Python, Scala, or Go - Experience managing Spark clusters and tuning Spark jobs - Active user of and/or contributor to open-source projects - Exceptional Skills - Familiar with Terraform, Docker, and Kubernetes - Familiar with managing data engineering infrastructure (Airflow, Kubernetes, etc) ## Why Viaduct - Contribute to the open-source ecosystem - Work with established experts in deep learning, time-series analytics, and convex optimization - Endless opportunities for technical learning and personal growth - Full health, vision, and dental benefits
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗