Principal Machine Learning Engineer
Core
Building proactive AI applications that organize users' lives by turning research and model capabilities into reliable, scalable production systems for long-running workflows and real-world task completion.
Role type
Principal Machine Learning Engineer (Production Systems & Architecture)
Builds
End-to-end ML systems, training/fine-tuning pipelines, evaluation systems, high-performance inference systems, and data pipelines for large models.
Domain
Artificial Intelligence / Large Language Models / Production Systems
Deliverable
production ML models
Required skills
End-to-end ML system ownership, large-model training and fine-tuning, evaluation system design, high-performance inference architecture, GPU-based system operations, data pipeline engineering, production infrastructure setup, technical trade-off analysis
Preferred skills
Experience shipping ML systems used by people, understanding large model failure modes, writing production-grade code
Technologies
Python, PyTorch, JAX, GPU-based training and inference systems
Responsibilities
Own end-to-end ML systems from data and training to evaluation, inference, and deployment; Build and evolve training and fine-tuning pipelines for large models; Design evaluation systems measuring capability, robustness, safety, and product performance; Architect high-performance inference systems optimizing latency, GPU utilization, memory, and cost; Build data pipelines for high-quality real-world and synthetic training data; Establish reliable production infrastructure for deploying and monitoring models; Partner with research and application engineering to improve product capabilities
Seniority
Principal, hands-on IC with leadership scope
