Principal Site Reliability Engineer (AIOps)
Core
Design, build, and operate reliable, secure cloud infrastructure for a large hybrid environment, focusing on automation, performance, and troubleshooting.
Role type
Principal Site Reliability Engineer (AIOps)
Builds
Production-ready, scalable, and reliable cloud services and automation frameworks
Domain
Cybersecurity, Cloud Infrastructure, AIOps
Deliverable
production ML models | infrastructure
Required skills
Linux administration, distributed systems troubleshooting, configuration management (Ansible, Terraform, Helm), Python, Golang, shell scripting, CI/CD pipelines, Kubernetes, Docker, GCP, AWS, monitoring and alerting orchestration
Preferred skills
GitLab, GitHub, Spinnaker, Datadog, Elasticsearch, Kafka, Hadoop, MySQL, Percona, MongoDB, Tensorflow
Responsibilities
Design and operate reliable, secure Cloud infrastructure; Develop tools and automation frameworks; Automate robust deployment of services; Orchestrate end-to-end monitoring and alerting; Lead root cause analysis of critical production issues; Mentor and champion SRE culture; Participate in design reviews
Seniority
Principal, hands-on IC with mentorship