Principal Cloud Infrastructure Engineer (Advanced Threat Protection)
Core
Design, build, and operate cloud-native infrastructure platforms powering applications across GCP, AWS, and global data centers, leveraging AI/ML to automate incident detection, root cause analysis, and remediation.
Role type
Senior Site Reliability Engineer (Cloud Infrastructure)
Builds
Intelligent, self-healing cloud infrastructure systems and automation tools for production applications
Domain
Cloud Infrastructure / Site Reliability Engineering / AI Operations (AIOps)
Deliverable
production ML models | infrastructure
Required skills
Multi-cloud infrastructure design (GCP, AWS, OCI), Infrastructure as Code (Terraform, Ansible), Python scripting, Linux distributed systems, CI/CD pipelines, HTTP/web servers/networking fundamentals, cloud compliance frameworks (FedRAMP, IL5)
Preferred skills
AI/ML application to operational workflows (AIOps, LLM-powered tooling), large database systems management (MySQL, PostgreSQL, Redis, BigQuery)
Technologies
GCP, AWS, OCI, Terraform, Ansible, Python, GitLab, Artifactory, MySQL, PostgreSQL, Redis, BigQuery
Responsibilities
Design and operate cloud infrastructure enabling reliable microservices deployment; Leverage AI/ML to automate incident detection and remediation; Build and integrate AI-powered tools into SRE workflows; Write automation code for provisioning infrastructure at massive scale; Develop self-healing systems for anomaly detection and corrective action; Lead root cause analysis of critical issues and build preventive automation; Mentor other SREs on infrastructure orchestration and AI-augmented operations
Seniority
Senior, hands-on IC