CareerPlanSign in

Staff Machine Learning Engineer

🌐 Remote💼 Full-time🗓 2026-07-23 → 2026-09-26

Core

Build infrastructure, tooling, and optimization systems to enable efficient training, evaluation, deployment, and operation of machine learning models at scale.

Role type

Staff Machine Learning Engineer (ML Infrastructure & Efficiency)

Builds

Scalable ML training and inference platforms, resource management systems, and performance monitoring tooling.

Domain

Internet / Machine Learning Infrastructure

Deliverable

infrastructure

Required skills

Distributed systems design, performance engineering, systems optimization, debugging and profiling, Python, at least one systems language (Go, C++, Rust, or Java)

Preferred skills

Large-scale recommendation or ranking systems, generative AI, foundation models, distributed training frameworks (PyTorch Distributed, Ray, Tensorflow, Spark), GPU architecture knowledge, cloud cost optimization

Technologies

PyTorch Distributed, Ray, Tensorflow, Spark, Go, C++, Rust, Java

Responsibilities

Design systems to improve efficiency of ML training and inference workloads; Develop tooling for debugging, profiling, optimizing, and monitoring model performance; Improve GPU and resource utilization through scheduling and caching; Partner with researchers to identify bottlenecks; Build benchmarking frameworks and performance dashboards; Optimize distributed training infrastructure and model serving architectures; Lead cross-functional initiatives for ML engineer productivity; Drive technical strategy for platform scalability and cost efficiency.

Seniority

Staff, hands-on IC with strategic leadership

Sourced via codingjobboard · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.