Site Reliability Engineer
Core
Ensure the reliability, scalability, and performance of critical intelligence platforms and infrastructure.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Scalable and reliable infrastructure on AWS, observability solutions, and automated provisioning pipelines.
Domain
Cybersecurity intelligence / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS, Linux, Terraform, Chef, Root Cause Analysis, Incident Management, Observability, Automation, Infrastructure as Code
Preferred skills
Kubernetes, RabbitMQ, Apache Kafka, MongoDB, Elasticsearch, OpenTelemetry, CI/CD pipelines, Microservices architecture
Technologies
AWS, Grafana, ELK Stack, Prometheus, Terraform, Chef, Kubernetes, RabbitMQ, Apache Kafka, MongoDB, Elasticsearch, OpenTelemetry
Responsibilities
Design and maintain scalable AWS infrastructure; Develop and manage observability solutions; Automate infrastructure provisioning; Perform Root Cause Analysis for outages; Participate in 24/7 on-call rotation; Collaborate with engineering teams on high availability; Proactively identify performance bottlenecks.
Seniority
Senior, hands-on IC