Site Reliability Engineer
Core
Build and maintain the reliability, scalability, and operational excellence of core systems including Kubernetes, AWS, CI/CD, and monitoring infrastructure.
Role type
Site Reliability Engineer (SRE)
Builds
Production environments, CI/CD pipelines, monitoring systems, and resilient infrastructure for software delivery.
Domain
Cybersecurity / Software Supply Chain Security / Cloud Infrastructure
Deliverable
production ML models | product features | dashboards & analysis | infrastructure
Required skills
Kubernetes, Linux, CI/CD, monitoring and observability, scripting (Bash/Python), networking fundamentals, Git workflows
Preferred skills
AWS, Terraform/Ansible, message queues (RabbitMQ), secrets management (Vault), CDN, incident response
Technologies
Kubernetes, Docker, GitLab CI, Jenkins, ArgoCD, Nginx, Apache, AWS, Terraform, Ansible, RabbitMQ, Vault, Fastly, Site24x7
Responsibilities
Develop pipelines for building and deploying scalable production environments; Define steps and processes for production management; Create and maintain monitoring and alerting systems; Estimate budget and hardware requirements; Manage day-to-day activities and task prioritization; Collaborate with engineers on technical design and infrastructure approach; Assess and recommend new technologies to enhance products.
Seniority
Mid-level, hands-on IC