Senior Machine Learning Systems Engineer
Core
Design and build infrastructure for large-scale machine learning systems, including graph ML platforms supporting billions of nodes/edges, optimizing training performance and GPU utilization.
Role type
Senior IC machine learning systems engineer (MLOps & distributed systems)
Builds
End-to-end MLOps patterns, graph ML platforms, scalable data processing pipelines, and ML infrastructure tooling
Domain
Cloud infrastructure + Graph Machine Learning
Deliverable
production ML models | infrastructure
Required skills
MLOps patterns, distributed systems, cloud data processing, GPU optimization, graph ML platforms, infrastructure-as-code, experiment tracking, model serving, Python, PyTorch/TensorFlow, Ray, Kubernetes
Preferred skills
Graph databases (Neo4j, JanusGraph, TigerGraph), Graph neural networks (PyTorch Geometric, Deep Graph Library)
Technologies
GCP BigQuery, Google Cloud Storage, Terraform, MLflow, Weights & Biases, Apache Beam, Apache Spark, Ray Data, Neo4j, JanusGraph, TigerGraph, PyTorch Geometric, Deep Graph Library
Responsibilities
Design end-to-end model lifecycle and MLOps workflows; Develop graph ML platforms; Optimize training performance and GPU costs; Architect pipelines for massive graph datasets; Administer experiment tracking and model registries; Build scalable, reliable infrastructure; Collaborate with users to reduce technical friction; Contribute to architectural decisions across cloud and ML systems.
Seniority
Senior, hands-on IC


