Staff Machine Learning Engineer - ML Frameworks
Core
Design, develop, and maintain robust AI/ML infrastructure solutions to support the training and deployment of large-scale generative AI models.
Role type
Staff Machine Learning Engineer (ML Infrastructure)
Builds
Distributed training frameworks, orchestration systems, and GPU resource management for Firefly generative AI models.
Domain
Generative AI, Cloud Infrastructure, Distributed Systems
Deliverable
production ML models
Required skills
Python, distributed training frameworks, GPU resource management, model orchestration and scheduling, PyTorch, Kubernetes, AWS, AutoML
Preferred skills
KubeFlow, MLFlow, Ray, SageMaker, PyTorch distributed, MPI, Megatron, Horovod
Responsibilities
Design and maintain AI/ML infrastructure for model training and deployment; Implement and improve distributed training frameworks leveraging GPUs; Optimize orchestration and scheduling for faster experimentation; Collaborate with data scientists to streamline training pipelines; Drive innovation in infrastructure practices.
Seniority
Staff, hands-on IC with strategic visibility


