Staff Machine Learning Engineer
Core
Build infrastructure, tooling, and optimization systems to enable efficient training, evaluation, deployment, and operation of machine learning models at scale.
Role type
Staff Machine Learning Engineer (ML Infrastructure & Efficiency)
Builds
Scalable ML training and inference platforms, resource management systems, and performance monitoring tooling.
Domain
Internet / Machine Learning Infrastructure
Deliverable
infrastructure
Required skills
Distributed systems design, performance engineering, systems optimization, debugging and profiling, Python, at least one systems language (Go, C++, Rust, or Java)
Preferred skills
Large-scale recommendation or ranking systems, generative AI, foundation models, distributed training frameworks (PyTorch Distributed, Ray, Tensorflow, Spark), GPU architecture knowledge, cloud cost optimization
Technologies
PyTorch Distributed, Ray, Tensorflow, Spark, Go, C++, Rust, Java
Responsibilities
Design systems to improve efficiency of ML training and inference workloads; Develop tooling for debugging, profiling, optimizing, and monitoring model performance; Improve GPU and resource utilization through scheduling and caching; Partner with researchers to identify bottlenecks; Build benchmarking frameworks and performance dashboards; Optimize distributed training infrastructure and model serving architectures; Lead cross-functional initiatives for ML engineer productivity; Drive technical strategy for platform scalability and cost efficiency.
Seniority
Staff, hands-on IC with strategic leadership
