Senior Software Engineer, Observability
Core
Designing, implementing, and maintaining robust distributed storage solutions and comprehensive observability platforms for a generative AI cloud infrastructure.
Role type
Senior Software Engineer, Observability
Builds
Scalable observability platforms (metrics, logs, traces), automated monitoring/alerting systems, and telemetry data pipelines for AI infrastructure.
Domain
Generative AI / Cloud Infrastructure / Observability
Deliverable
infrastructure
Required skills
Observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry), Go/Python programming, Infrastructure-as-Code (Terraform, Ansible, Helm), Distributed systems design, Containerization (Docker), Orchestration (Kubernetes), Time-series databases, Microservices architecture, CI/CD pipelines, GitOps workflows, Database management (PostgreSQL, MongoDB, Redis).
Preferred skills
AI/ML infrastructure monitoring, GPU cluster monitoring, High-frequency low-latency systems monitoring, Chaos engineering, Reliability testing, Open-source observability contributions, Security monitoring frameworks.
Technologies
Prometheus, Grafana, ClickHouse, ClickStack, OpenTelemetry, Go, Python, Terraform, Ansible, Helm, Docker, Kubernetes, PostgreSQL, MongoDB, Redis, AWS, GCP, Azure.
Responsibilities
Design and implement scalable observability platforms including telemetry data pipelines and log aggregation workflows; Develop automated monitoring, alerting, and anomaly detection systems with SLIs/SLOs and predictive analytics; Build and deploy custom observability tools and infrastructure-as-code; Collaborate with engineering teams to enhance distributed tracing and lead incident response; Define observability best practices.
Seniority
Senior, hands-on IC