CareerPlanSign in

Senior ML Data Processing Developer

Montreal💼 Full-time🗓 2026-07-02 → 2026-09-26

Core

Design and engineer end-to-end data pipelines to transform raw web-scale data into high-signal training datasets for the Scientist AI.

Role type

Senior ML Data Processing Developer

Builds

Production-grade data processing pipelines, quality scoring toolchains, and internal dataset exploration interfaces for the Scientist AI.

Domain

AI Safety / NLP / Data Engineering

Deliverable

production ML models

Required skills

Python, distributed processing frameworks (Spark, Ray, Flink), pipeline orchestration (Airflow, Prefect, Dagster), data privacy implementation, content-safety filtering, evaluation-contamination prevention, large-scale unstructured text dataset handling, LLM-as-a-judge evaluators, metadata extraction, data licensing workflows

Preferred skills

ML model training/fine-tuning for data-quality tasks, LLM inference optimization (vLLM, SGLang), containerized deployment (Docker, Kubernetes), infrastructure-as-code, ML experiment tracking (Weights and Biases), open-source contributions

Technologies

Spark, Ray, Flink, Airflow, Prefect, Dagster, vLLM, SGLang, Docker, Kubernetes, Weights and Biases

Responsibilities

Partner with Research to define, build, automate, and scale data pipelines; Build and maintain pipelines including deduplication, quality scoring, heuristic filtering, toxicity removal, PII scrubbing, and metadata extraction; Develop and refine scoring/filtering toolchains using heuristics, LLM evaluators, and ML classifiers; Instrument pipelines with data-quality monitoring, guardrails, and alerting; Identify and acquire large-scale text corpora meeting requirements; Design and maintain leakage detection mechanisms; Build internal tooling for researchers to explore and query datasets.

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.