Senior Software Engineer- Site Reliability
Core
Design, build, and manage AWS infrastructure and Kubernetes clusters to ensure the reliability, scalability, and security of critical financial applications and customer-facing services.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Production services supporting critical financial applications on AWS EKS
Domain
Financial services, Cloud Infrastructure, Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
Python development for automation, AWS infrastructure management, Kubernetes cluster operations, Terraform (Infrastructure as Code), networking fundamentals, incident management, root cause analysis, observability tuning, CI/CD pipeline management, database query optimization
Preferred skills
Experience with Kafka or other messaging systems, familiarity with Helm, contributions to open-source infrastructure projects, experience in regulated environments
Technologies
AWS, Amazon EKS, Terraform, Jsonnet, Kubernetes, Helm, Prometheus, Grafana, PostgreSQL, Linux (Ubuntu)
Responsibilities
Own and operate production services with a focus on availability and performance; Design and manage AWS infrastructure including EKS clusters; Provision infrastructure using Terraform; Deploy, scale, and troubleshoot applications on Kubernetes; Build automation frameworks to reduce operational toil; Monitor system health and tune alerts and dashboards; Lead incident response and drive root cause analysis; Collaborate with security teams to maintain security posture; Optimize infrastructure cost and resource utilization
Seniority
Senior, hands-on IC