Machine Learning Systems Engineer, Networking
Core
Design and implement real-time ML algorithms for anomaly detection and predictive analytics on high-volume GPU fleet telemetry within strict resource constraints.
Role type
Senior IC Machine Learning Systems Engineer
Builds
Real-time streaming pipelines for AI Data Center AIOps platform
Domain
Cloud Infrastructure / GPU Fleet Management / Time-Series Telemetry
Deliverable
production ML models
Required skills
Go, C/C++, Rust, Scala, time-series databases, streaming data architectures, anomaly detection, predictive analytics, algorithm optimization, statistical modeling, linear algebra
Preferred skills
Kafka-based streaming pipelines, real-time feature engineering, research experience translating ML literature to production
Technologies
Go, C/C++, Rust, Scala, Kafka, time-series databases
Responsibilities
Implement production ML algorithms optimized for real-time streaming pipelines, design and develop new ML algorithms for anomaly detection and health scoring, build and maintain end-to-end ML pipelines from ingestion to inference, partner with Data Science team on algorithm design and prototype evaluation
Seniority
Senior, hands-on IC