Staff Site Reliability Engineer
Core
Senior technical leadership role owning multi-week SRE projects, architectural decisions, and team mentorship for an AI-powered energy intelligence platform.
Role type
Staff Site Reliability Engineer (L4)
Builds
AWS-based infrastructure, Kubernetes clusters, CI/CD pipelines, and observability stacks for enterprise energy management.
Domain
Energy & Utilities / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
AWS architecture, Terraform, Kubernetes operations, CI/CD pipeline design, observability stack management, FinOps, database reliability, security posture management, incident management, technical mentorship, autonomous execution
Preferred skills
FinOps practices, secrets management, event-driven architectures, AI-enabled tooling, data warehouse operations, workflow automation platforms, industry certifications
Technologies
AWS (EKS, VPC, RDS, IAM, CloudWatch, CloudTrail, GuardDuty, Lambda, SQS, S3, CloudFront), Terraform, CloudFormation, Jenkins, GitHub Actions, ArgoCD, FluxCD, Prometheus, Grafana, Datadog, MySQL, PostgreSQL, Vault, AWS Secrets Manager
Responsibilities
Own and deliver SRE projects end-to-end, serve as a technical anchor for the India SRE team, design and implement infrastructure solutions, lead Kubernetes operations, evolve CI/CD pipelines, drive observability stack enhancements, identify and execute FinOps initiatives, manage database reliability, strengthen security posture, troubleshoot complex cross-cutting production issues, write documentation, collaborate with US-based SRE leadership, participate in on-call rotations
Seniority
Staff, technical leadership & mentorship