Member of Technical Staff, Synthetic Data
Core
Design and build scalable inference pipelines and synthetic data curation methods to improve token efficiency and model quality for frontier language models.
Role type
Senior Machine Learning Engineer (Synthetic Data)
Builds
Scalable inference pipelines running on large GPU clusters and synthetic data pipelines for advanced language models
Domain
Natural Language Processing (NLP) and Large Language Models (LLMs)
Deliverable
production ML models
Required skills
Python, data pipeline development, Apache Spark, Apache Beam, Pandas, LLMs, vLLM, TensorRT, large-scale dataset handling
Preferred skills
Research and engineering bridge, data ablation techniques, data mixture experimentation, top-tier conference publications
Technologies
Python, Apache Spark, Apache Beam, Pandas, vLLM, TensorRT
Responsibilities
Design and build scalable inference pipelines on large GPU clusters; Conduct data ablations to assess data quality and experiment with data mixtures; Research and implement innovative synthetic data curation methods; Collaborate with cross-functional teams to ensure data pipelines meet model demands
Seniority
Senior, hands-on IC