Senior Software Engineer- Site Reliability
Core
Ensure reliability, scalability, and security of business-critical internal systems and external customer-facing services through hands-on infrastructure and software engineering.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Resilient AWS infrastructure, EKS clusters, and automation frameworks for financial applications.
Domain
Financial services / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python development, AWS (IAM, networking), Kubernetes (EKS), Terraform, incident management, root cause analysis, observability, CI/CD pipelines, distributed systems design
Preferred skills
Jsonnet, Helm, Kafka, PostgreSQL, Linux (Ubuntu), security posture management, cost optimization
Technologies
AWS, Amazon EKS, Terraform, Jsonnet, Kubernetes, Helm, Prometheus, Grafana, PostgreSQL, Kafka, Linux
Responsibilities
Design, build, and manage AWS infrastructure including EKS-based clusters; Provision and manage infrastructure using Terraform; Deploy, scale, and troubleshoot applications on Kubernetes; Build and maintain Python-based automation frameworks; Monitor system health and tune alerts/dashboards; Lead incident response and drive root cause analysis; Collaborate with security teams to maintain security posture; Optimize infrastructure cost and resource utilization.
Seniority
Senior, hands-on IC