SRE Engineer - CyberSecurity
Core
Implement, monitor, and maintain software systems to ensure high availability, performance, and reliability through automation and incident management.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Scalable enterprise systems, intelligent automation frameworks, and digital transformation platforms
Domain
IT Consulting, Cloud Infrastructure, DevOps
Deliverable
production ML models | product features | infrastructure
Required skills
Linux/Unix systems administration, containerization and orchestration (Docker, Kubernetes), automation tools (Puppet, Chef, Ansible), scripting (Python, Bash, Perl), cloud platforms (AWS, GCP), chaos engineering, load testing, capacity management
Preferred skills
Agile development environments, third-party platform competencies
Technologies
Docker, Kubernetes, Puppet, Chef, Ansible, Python, Bash, Perl, AWS, GCP
Responsibilities
Implement, monitor, and maintain software systems for high availability; Automate repetitive tasks and processes; Develop and maintain monitoring and analysis tools; Collaborate with development teams on reliable and scalable service deployment; Respond to and resolve incidents and outages; Conduct post-incident analysis and documentation; Participate in on-call rotation for 24x7 support; Write SRE best practices and business rules; Support usage/cost allocation audits; Support release roadmap planning; Conduct chaos engineering activities; Perform load testing and capacity management tasks.
Seniority
Mid-level, hands-on IC