Senior Data & MLOps Engineer
Core
Design and scale the infrastructure supporting the GPU Intelligence Platform, building pipelines for data, features, model training, and delivering insights for system health and optimization.
Role type
Senior IC Data & MLOps Engineer
Builds
Scalable distributed services, data ingestion pipelines, feature processing systems, and production ML models for GPU fleet reliability.
Domain
AI Infrastructure / GPU Telemetry / Distributed Systems
Deliverable
production ML models | infrastructure
Required skills
Data engineering, distributed systems, MLOps, high-throughput streaming pipelines, time-series feature pipelines, feature stores, ML model deployment (versioning/monitoring/rollback/drift), microservices architecture, Python, systems languages (Go/Rust/C++)
Preferred skills
GPU fleet management, hyperscale infrastructure, anomaly detection, failure prediction, distributed straggler detection, agentic/LLM-powered reasoning systems, reliability engineering (SRE)
Technologies
Kafka, Pulsar, Kinesis, Kubernetes, NCCL, PyTorch Distributed, Spark, Ray, Slurm, NVML, DCGM
Responsibilities
Design and implement scalable data ingestion pipelines; Build feature processing and baseline computation systems; Productionize models for prediction and detection; Develop and operate low-latency service and robust offline workflows; Architect horizontally scalable services with clear separation between components; Implement monitoring and feedback loops for continuous model and signal improvement; Collaborate with Platform teams to integrate operational signals into monitoring and diagnostics
Seniority
Senior, hands-on IC