Staff Machine Learning Systems Engineer
Core
Designing and building large-scale ML infrastructure and platforms to support model training, deployment, and data processing for Reddit's recommendation and content discovery systems.
Role type
Staff Machine Learning Systems Engineer (Infrastructure)
Builds
End-to-end MLOps platforms, graph ML codebases, and massive graph data structures (billions of nodes/edges)
Domain
Social Media / Distributed Machine Learning Systems
Deliverable
production ML models
Required skills
MLOps architecture, distributed training optimization, graph database management, cloud infrastructure (GCP), MLOps tool integration, Python, PyTorch, TensorFlow, Ray, Kubernetes
Preferred skills
Graph neural networks (GNNs), graph ML frameworks (PyTorch Geometric, Deep Graph Library), graph databases (Neo4j, JanusGraph, TigerGraph)
Technologies
GCP BigQuery, Google Cloud Storage, Terraform, Apache Beam, Apache Spark, Ray Data, MLflow, Wandb, Neo4j, JanusGraph, TigerGraph, PyTorch Geometric, Deep Graph Library
Responsibilities
Design end-to-end model lifecycle patterns (MLOps) to boost development velocity; Develop and support a graph ML codebase and platform; Collaborate on performance tuning for model training and GPU efficiency; Optimize batch data processing within data warehouses; Architect pipelines for massive graph data structures
Seniority
Staff, hands-on IC with strategic scope