Lead Machine Learning Engineering, (Hybrid)
Core
Build and improve scalable data pipelines, human-in-the-loop labeling workflows, and synthetic data generation systems to create high-quality training and evaluation datasets for Large Language Models (LLMs) and AI agents.
Role type
Lead Machine Learning Engineer (Data Engineering focus)
Builds
Scalable data pipelines, labeling workflows, synthetic datasets, and automated evaluation systems for LLMs
Domain
Networking, Generative AI, Large Language Models (LLMs)
Deliverable
production ML models
Required skills
Python, C++, Go, PyTorch, TensorFlow, distributed data processing (Spark, Ray, Beam), dataset curation, human-in-the-loop labeling, synthetic data generation, bias mitigation, model evaluation
Preferred skills
LLM lifecycle management (SFT, RLHF), model-assisted labeling, distributed data architecture, research-engineering mindset
Technologies
PyTorch, TensorFlow, Spark, Ray, Beam, Python, C++, Go
Responsibilities
Design and maintain robust data pipelines for ML/LLM development; Architect human-in-the-loop labeling workflows; Develop strategies for synthetic data generation and validation; Leverage LLMs to automate data generation and evaluation; Establish systems to measure and mitigate dataset failure modes; Collaborate with researchers to define dataset requirements; Provide technical direction on infrastructure and mentor the team
Seniority
Senior, hands-on IC with leadership responsibilities