Senior Software Engineer- Site Reliability
Core
Ensure reliability, scalability, and security of business-critical internal systems and external customer-facing services through infrastructure ownership and automation.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Resilient AWS infrastructure, EKS clusters, and automation frameworks for financial applications.
Domain
Financial services / Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python development, AWS (EKS, IAM, networking), Kubernetes, Terraform, incident management, root cause analysis, observability, CI/CD pipelines, Linux.
Preferred skills
Jsonnet, Helm, Kafka, PostgreSQL, security posture management, cost optimization.
Technologies
AWS, Amazon EKS, Terraform, Jsonnet, Kubernetes, Helm, Prometheus, Grafana, PostgreSQL, Kafka, Linux (Ubuntu)
Responsibilities
Own and operate production services supporting critical financial applications; Design, build, and manage AWS infrastructure including EKS-based clusters; Provision and manage infrastructure using Terraform; Deploy, scale, and troubleshoot applications running on Kubernetes; Build and maintain automation frameworks and tooling Python based; Monitor system health using metrics, logs, and alerts; Troubleshoot complex issues spanning clusters, networking, certificates, deployments, and application behavior; Manage certificate lifecycle and expiration; Collaborate with InfoSec, Vulnerability Management, and Network Security teams; Participate in on call and lead incident response; Identify architectural anti-patterns and drive improvements; Establish and enforce production readiness standards; Optimize infrastructure cost and resource utilization.
Seniority
Senior, hands-on IC