Staff Machine Learning Systems Engineer
Core
Design and scale cloud-based infrastructure for large-scale machine learning systems, specifically focusing on graph ML platforms handling billions of nodes and edges.
Role type
Staff Machine Learning Systems Engineer (Infrastructure & MLOps)
Builds
Scalable graph machine learning platforms, MLOps pipelines, and distributed training environments
Domain
Cloud Infrastructure, Distributed Machine Learning, Graph Data Systems
Deliverable
production ML models | infrastructure
Required skills
MLOps pattern design, distributed ML systems, cloud infrastructure architecture, graph data processing, performance optimization, infrastructure-as-code, Python, Kubernetes, Ray, GPU resource management
Preferred skills
Graph databases (Neo4j, JanusGraph), Graph Neural Networks (PyTorch Geometric, DGL), Google Cloud Platform (BigQuery, GCS)
Technologies
Terraform, MLflow, Weights & Biases, Apache Beam, Apache Spark, Ray Data, Kubernetes, PyTorch, TensorFlow, Neo4j, JanusGraph, TigerGraph, BigQuery, Google Cloud Storage
Responsibilities
Design end-to-end model lifecycle patterns and MLOps capabilities; Lead zero-to-one development of graph ML infrastructure; Optimize model training performance and GPU utilization; Architect pipelines for massive graph data structures; Build and evolve cloud-based ML platforms; Administer MLOps tools for experiment tracking and model serving; Establish engineering patterns to improve platform reliability and developer productivity; Provide technical leadership for long-term ML system architecture
Seniority
Staff, hands-on IC with significant technical ownership and cross-functional leadership

