Principal Group Engineering Manager
Core
Build and evolve distributed services underpinning massive training runs, improve iteration loops for researchers, and apply performance optimization techniques to impact product quality and operational excellence.
Role type
Principal Group Engineering Manager (ML Systems & AI Infrastructure)
Builds
Distributed services, AI infrastructure, compute orchestration, and AI-Native development practices.
Domain
Machine Learning, AI Infrastructure, Cloud Computing, Compute Orchestration
Deliverable
production ML models | infrastructure
Required skills
Distributed systems design, ML systems architecture, AI infrastructure, compute orchestration, performance optimization, code performance tuning, system architecture, technical leadership, SDLC governance, Responsible AI practices, GenAI tooling adoption, cloud software development, debugging large codebases
Preferred skills
Reinforcement learning, post-training, container workloads, C/C++/C#/Java/JavaScript/Python
Technologies
GenAI tools (e.g., GitHub CoPilot), container workloads, cloud platforms
Responsibilities
Lead technical strategy for AI-Native development and Responsible AI practices across the SDLC; oversee coding standards, performance, and security for extensible and maintainable code; own complex system architecture design and evaluation; drive engineering excellence through best practices and GenAI tooling adoption; coordinate project plans and release schedules across multiple groups.
Seniority
Principal, hands-on IC with team leadership