Staff Software Engineer
Core
Design, build, and maintain scalable, secure, and reliable cloud infrastructure while implementing Site Reliability Engineering (SRE) practices to ensure system performance and availability.
Role type
Staff Cloud Engineer (SRE)
Builds
Cloud-native infrastructure, CI/CD pipelines, and automated deployment systems for the Harness AI Software Delivery Platform
Domain
Cloud Infrastructure & Site Reliability Engineering (SRE)
Deliverable
production ML models | infrastructure
Required skills
Cloud platforms (AWS, GCP, Azure), Infrastructure-as-Code (Terraform, CloudFormation), Kubernetes (K8s), Helm, CI/CD pipelines, Python, Go, Monitoring/Observability (Prometheus, Grafana), Security & Compliance frameworks
Preferred skills
AI-OPS SRE practices, Error budgets, SLIs/SLOs, Incident management, Ansible, Chef, Puppet, Jenkins, GitLab CI, CircleCI
Technologies
GCP, AWS, Azure, Terraform, CloudFormation, Kubernetes, Helm, Prometheus, Grafana, Jenkins, GitLab CI, CircleCI, Ansible, Chef, Puppet
Responsibilities
Design and implement scalable cloud infrastructure using IaC; Develop and maintain monitoring, logging, and alerting systems; Perform capacity planning and demand forecasting; Deploy and manage applications using Kubernetes and Helm; Design and implement CI/CD pipelines; Ensure cloud infrastructure meets security and compliance standards; Mentor junior engineers
Seniority
Staff, hands-on IC with mentorship