Observability, Data Scientist
Core
Transform large-scale infrastructure and hardware-level telemetry into actionable insights, predictive intelligence, and automated decision systems for AI compute stacks.
Role type
Data Scientist specializing in infrastructure observability and hardware telemetry
Builds
Automated remediation strategies, predictive failure models, and closed-loop optimization systems for data center and chip-level infrastructure
Domain
Semiconductor hardware, data center infrastructure, AI compute stack
Deliverable
production ML models | dashboards & analysis
Required skills
time-series data analysis, anomaly detection, predictive modeling, signal processing, Python (NumPy, Pandas, SciPy, ML frameworks), data visualization, distributed systems understanding
Preferred skills
data center infrastructure experience, hyperscaler observability solutions, root cause analysis in distributed systems, model deployment, system architecture knowledge, real-time analytics frameworks
Technologies
Prometheus, InfluxDB, Grafana, Superset, Tableau, Plotly, Kafka, Flink, Spark