Staff Infrastructure Engineer — Observability
Core
Architect and implement robust, scalable telemetry platforms to provide real-time global visibility and actionable insights for SentinelOne's AI-native security platform.
Role type
Staff Infrastructure Engineer (Observability)
Builds
Core observability stack (Grafana, Prometheus, Thanos/Mimir/Cortex, OpenTelemetry) and high-volume data ingestion pipelines for a global security platform.
Domain
Cybersecurity / Cloud Infrastructure / Observability
Deliverable
production ML models | infrastructure
Required skills
Architecting enterprise-grade observability stacks, cloud-native infrastructure design, Kubernetes management, IaC (Terraform/Ansible), distributed systems optimization, technical leadership, incident management.
Preferred skills
GoLang programming, high-security compliance frameworks (FedRAMP), on-premises/hybrid Kubernetes operations, advanced CI/CD pipeline design.
Technologies
Grafana, Prometheus, Thanos, Mimir, Cortex, OpenTelemetry, AWS, GCP, Kubernetes (EKS, GKE), Terraform, Ansible, GitHub Actions.
Responsibilities
Architect and implement scalable telemetry platforms; act as SME for the core observability stack; partner with engineering teams to define platform requirements; take end-to-end ownership of critical features; drive operational efficiency and cost-optimization on AWS/GCP; build automation and self-service tooling; ensure compliance in high-security environments; mentor engineers and lead technical design reviews; resolve complex production incidents.
Seniority
Staff, strategic architect & hands-on IC