Senior AI and HPC Observability Engineer
Core
Design and scale next-generation observability and telemetry platforms for high-volume metrics, logs, and traces across distributed AI/HPC environments.
Role type
Senior IC backend/distributed systems engineer (observability)
Builds
High-throughput telemetry pipelines, backend services, and monitoring frameworks for AI superclusters
Domain
High-performance computing (HPC) and Artificial Intelligence (AI)
Deliverable
production ML models | infrastructure
Required skills
Python, Go, Java, distributed systems design, time-series data systems, streaming technologies, Kubernetes, fault-tolerant system design
Preferred skills
OpenTelemetry, Prometheus, Kafka, Spark, Flink, GPU workload monitoring, statistical anomaly detection
Technologies
OpenTelemetry, Prometheus, Kafka, Spark, Flink, Kubernetes, PromQL
Responsibilities
Design and scale observability platforms handling high-volume metrics, logs, and traces; Build high-performance backend services for telemetry ingestion, processing, and routing; Develop and extend OpenTelemetry collectors, processors, exporters, and instrumentation libraries; Build and optimize metrics pipelines using large-scale time-series storage systems; Design and operate real-time and batch telemetry pipelines using streaming and distributed data technologies; Improve platform reliability, performance, and cost efficiency through tuning, capacity planning, and system optimization; Develop monitoring, alerting, and service reliability frameworks to ensure platform health and performance
Seniority
Senior, hands-on IC