💼 Full-time🗓 2026-07-26
Rewrite
## About the role
As our Data Ops Engineer, you'll:
- Manage and version high-quality datasets used for training and evaluating speech/AI models.
- Build and maintain lightweight ETL pipelines and scripts for data ingestion, processing, and automation.
- Ensure data quality through validation checks, consistency reviews, and monitoring.
- Run recurring benchmarks and maintain simple dashboards/reports to track model and data performance.
- Collaborate with product, ML, and engineering teams to support new data needs and improve workflows.
## Requirements
- 1–2 years in a data engineering/ops, ML ops‑adjacent, or analytics engineering role
- Python (Pandas/NumPy), Git, and Linux comfort for day‑to‑day scripting
- SQL fundamentals for joins/filters/aggregates
- Proven attention to detail (you notice when 2% of files go missing or WER shifts by 0.3%)
## Nice to have
- Worked with ASR/NLP data (transcripts, captions, subtitles)
- Familiarity with S3/object storage and experiment tracking (W&B/MLflow)
- Exposure to ASR/diarization tools (e.g., WhisperX, AssemblyAI, pyannote)
- Basics of SFT/RLHF concepts and evaluation metrics (WER/CER/DER)
## Tooling you'll touch
- Python + Pandas/NumPy, CLI scripts
- Git, Linux, simple packaging/virtualenvs
- SQL, S3/object storage
- (Optional) W&B/MLflow, Prefect/Airflow; ASR/diarization libs
## Why this role is interesting
- High leverage: Your work directly speeds up research and improves product quality.
- Broad exposure: Audio, text, diarization, evaluation, and RL datasets—without needing to be a full‑time modeler.
- Growth path: Data Ops → ML Engineering as you pick up more modeling/eval depth.
## Work style
- Hybrid, Bangalore: on-site 2–3 days/week for tight loops with research & product.
- Pragmatic, script‑first workflows with strong docs and versioning.
- Small team, high ownership, quick decisions.
## About the company
Scribie is an AI-powered, Human Verified audio and video transcription service, trusted globally since 2008. We specialize in delivering accurate and reliable transcription solutions by blending advanced AI technology with human expertise. Headquartered in the US, we operate with a hybrid model in our Bangalore office, combining the flexibility of remote work with the collaboration of in-person engagement. This approach offers our team both autonomy and growth opportunities in a dynamic and supportive environment.
## Experience
- 1–2 years
## Location
- Bangalore (hybrid)
## Type
- Full‑time
## How to apply
Send your resume/LinkedIn and (if you have them) links to GitHub/Kaggle/notebooks where you've cleaned or evaluated real data.
Optional mini‑signal: share a brief note on a data quality check you've used before that caught a subtle bug.
## Interview process
- Intro call (30 min)
- Practical take‑home (2–3 hrs)
- Onsite/virtual deep dive (60–90 min)
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.