Staff Machine Learning Engineer, ML Efficiency
Core
Building infrastructure, tooling, and optimization systems to enable efficient training, evaluation, deployment, and operation of machine learning models at scale.
Role type
Staff Machine Learning Engineer (ML Efficiency/Infrastructure)
Builds
Scalable ML training and inference infrastructure, optimization tooling, benchmarking frameworks, and performance dashboards.
Domain
Internet / Machine Learning Infrastructure / Systems Engineering
Deliverable
infrastructure
Required skills
Python, distributed systems, performance engineering, systems optimization, debugging and profiling, ML infrastructure, training systems, model serving platforms
Preferred skills
large-scale recommendation/ranking/generative AI/foundation models, distributed training frameworks (PyTorch Distributed, Ray, Tensorflow, Spark), GPU architecture knowledge, cloud cost optimization, real-time ML inference
Technologies
Python, Go, C++, Rust, Java, PyTorch Distributed, Ray, Tensorflow, Spark
Responsibilities
Design and build systems to improve efficiency of ML training and inference workloads; Develop tooling for debugging, profiling, optimizing, and monitoring model performance; Improve GPU and resource utilization through scheduling and caching; Partner with researchers to identify bottlenecks; Build benchmarking frameworks and performance dashboards; Optimize distributed training infrastructure and model serving architectures; Lead cross-functional initiatives to improve ML engineer productivity; Drive technical strategy for platform scalability and cost efficiency.
Seniority
Staff, hands-on IC with strategic impact