Site Reliability Engineer
Core
Ensuring the reliability, scalability, and performance of critical systems for a global intelligence platform.
Role type
Site Reliability Engineer (SRE)
Builds
Scalable and reliable infrastructure on AWS
Domain
Cyber threat intelligence / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS, Linux, Terraform, Chef, Observability (Grafana, ELK, Prometheus), Root Cause Analysis, Incident Management, Automation, System Architecture
Preferred skills
Kubernetes, Message Brokers (RabbitMQ, Kafka), NoSQL (MongoDB, Elasticsearch), OpenTelemetry, CI/CD pipelines, Microservices
Technologies
AWS, Terraform, Chef, Grafana, ELK, Prometheus, Kubernetes, RabbitMQ, Kafka, MongoDB, Elasticsearch, OpenTelemetry
Responsibilities
Ensure performance, capacity, scalability, reliability, and security of the platform; Perform Root Cause Analysis for outages; Design and maintain scalable AWS infrastructure; Develop and manage observability solutions; Automate infrastructure provisioning; Participate in 24/7 on-call rotation; Collaborate with engineering teams on high availability design.
Seniority
Mid-Senior, hands-on IC