Data Ops Engineer
Core
Manage and version high-quality datasets for training and evaluating speech/AI models, building ETL pipelines and ensuring data quality.
Role type
Junior data operations engineer (ML ops-adjacent)
Builds
Lightweight ETL pipelines, data validation scripts, and performance dashboards for speech/AI models
Domain
AI/ML data infrastructure for audio and video transcription
Deliverable
production ML models
Required skills
Python (Pandas/NumPy), Git, Linux, SQL, data versioning, data quality validation
Preferred skills
ASR/NLP data handling, S3/object storage, experiment tracking (W&B/MLflow), ASR/diarization tools, SFT/RLHF concepts
Technologies
Python, Pandas, NumPy, Git, Linux, SQL, S3, W&B, MLflow, Prefect, Airflow
Responsibilities
Manage and version datasets for model training/evaluation, build and maintain ETL pipelines, ensure data quality via validation checks, run benchmarks and maintain dashboards, collaborate with product/ML teams
Seniority
Junior, hands-on IC
