Principal Machine Learning Infrastructure Engineer, Ads & Discovery
Core
Building large-scale ML infrastructure to power ads and discovery systems for hundreds of millions of users, focusing on model training, serving, and optimization.
Role type
Principal Machine Learning Infrastructure Engineer
Builds
Scalable production-ready ML systems for recommendation, search, and agentic applications
Domain
Internet / Machine Learning Infrastructure
Deliverable
production ML models
Required skills
Large-scale ML system design, distributed training, low-latency inference, GPU optimization, model quantization/pruning, custom kernel development, strategic planning, architecture ownership
Preferred skills
Experience with transformer architectures, LLMs, generative rankers, FSDP, vLLM, SGLang, CUDA
Technologies
FSDP, vLLM, SGLang, CUDA, distributed training frameworks, inference engines, GPU kernels
Responsibilities
Co-design models and systems at the intersection of architecture and infrastructure; investigate tradeoffs from data pipelines to production serving; lead strategic planning and roadmap execution; establish engineering best practices for scalability and cost-effectiveness; collaborate with data scientists and product teams to operate robust ML platforms; stay abreast of industry trends in ML and infrastructure
Seniority
Principal, strategy & mentorship
