Principal Machine Learning Engineer
Core
Build and deploy production-grade machine learning systems for an AI-native communication platform, focusing on translating research into scalable solutions for reliable AI workflows and task execution.
Role type
Principal Machine Learning Engineer (LLM/Systems)
Builds
Scalable inference systems, training pipelines, evaluation frameworks, and deployment infrastructure for an AI assistant.
Domain
AI/ML, Large Language Models, Production Systems
Deliverable
production ML models
Required skills
Deep learning, transformer architectures, model fine-tuning (LoRA, QLoRA, SFT, DPO), distributed training, GPU optimization, software engineering, system design
Preferred skills
LLM inference frameworks (vLLM, TensorRT-LLM), open-source contributions, compiler technologies, RLHF pipelines, multimodal/diffusion models, large-scale data processing
Technologies
PyTorch, JAX, DeepSpeed, FSDP, Megatron, ZeRO, Ray, vLLM, TensorRT-LLM, FasterTransformer
Responsibilities
Build end-to-end ML pipelines (data processing, training, evaluation, inference, deployment); Fine-tune and adapt models using modern techniques; Design and operate scalable inference systems; Develop data pipelines for synthetic and real-world datasets; Build evaluation frameworks for model performance and safety; Optimize production deployments for latency and cost; Collaborate on integrating ML systems into backend and mobile applications.
Seniority
Principal, hands-on IC with strategic ownership