Principal Machine Learning Engineer
Core
Build and deploy production-grade machine learning systems for an AI-native communication platform, translating research into scalable solutions for real-world constraints.
Role type
Principal Machine Learning Engineer (LLM/Systems)
Builds
Production ML pipelines, inference systems, evaluation frameworks, and deployment infrastructure for an AI assistant platform.
Domain
AI / Large Language Models / Distributed Systems
Deliverable
production ML models
Required skills
Deep learning, transformer architectures, model training and fine-tuning, distributed training, software engineering, GPU optimization, system design
Preferred skills
LLM inference frameworks, open-source contributions, compiler technologies, RLHF pipelines, multimodal/diffusion models, large-scale data processing
Technologies
PyTorch, JAX, DeepSpeed, FSDP, Megatron, ZeRO, Ray, vLLM, TensorRT-LLM, FasterTransformer, Apache Arrow, Spark
Responsibilities
Build end-to-end ML pipelines (data processing, training, evaluation, inference, deployment); Fine-tune models using LoRA, QLoRA, SFT, DPO, and distillation; Design scalable inference systems balancing latency and cost; Develop data pipelines for synthetic and real-world datasets; Build evaluation frameworks for performance, robustness, safety, and bias; Optimize production deployments via GPU optimization and scaling; Integrate ML systems into backend, mobile, and desktop applications.
Seniority
Principal, hands-on IC with strategic ownership
