Sr. SDE, Edge AI ML Platform, Edge AI and Science
Core
Lead the architecture and delivery of core ML platform capabilities for training, optimizing, evaluating, and deploying generative AI models on devices and in the cloud.
Role type
Senior Software Development Engineer (ML Platform Architecture)
Builds
Distributed ML platform services and libraries for model ingestion, optimization, training, evaluation, packaging, and deployment
Domain
Edge AI, Generative AI, High-Performance Computing, Distributed Systems
Deliverable
production ML models
Required skills
distributed systems design, high-performance computing, system architecture, technical leadership, mentoring, cross-team collaboration
Preferred skills
distributed ML training/inference platforms, containerization, Kubernetes, AWS infrastructure, model compression, quantization, knowledge distillation, model compilation, edge deployment, extensible platform APIs
Technologies
PyTorch, TensorFlow, JAX, NeMo, Megatron, Kubernetes, AWS
Responsibilities
Lead design and delivery of distributed ML platform services; Define stable APIs and architecture boundaries; Design distributed training capabilities across parallelism types; Scale workflows on multi-node GPU clusters; Develop infrastructure for model optimization techniques; Build evaluation and artifact workflows; Build automated validation, CI/CD, and observability mechanisms; Profile and optimize end-to-end system performance; Establish operational mechanisms for production services; Partner with model, compiler, runtime, hardware, security, and infrastructure teams; Write technical designs and build consensus; Mentor engineers and improve code review practices
Seniority
Senior, hands-on IC with technical leadership