Principal Infrastructure Engineer
Core
Own technical architecture and evolution of core infrastructure, building capacity models, running load/stress tests, and designing resilient multi-region AWS and Kubernetes platforms.
Role type
Principal Infrastructure Engineer
Builds
Resilient cloud infrastructure (AWS, Kubernetes, Aurora RDS), observability tooling, and AI-assisted operational automation.
Domain
Cloud Infrastructure / Site Reliability Engineering
Deliverable
production ML models | infrastructure
Required skills
AWS architecture (IAM, networking, multi-AZ), Kubernetes cluster lifecycle and autoscaling, Aurora RDS optimization (MySQL/Postgres), Golang/Python, Terraform, Linux systems fundamentals, distributed systems failure modes, observability (metrics/logs/traces), disaster recovery planning, capacity modeling, load testing, CI/CD automation.
Preferred skills
Fintech/payments/banking experience, chaos engineering, Prometheus/Grafana/Loki/Tempo, AI-assisted tooling for incident investigation and toil reduction.
Responsibilities
Design resilient multi-region architectures; build and operate Kubernetes platforms; scale and optimize Aurora RDS; implement reliability engineering practices (error budgets, backpressure); participate in on-call rotations and incident recovery; design disaster recovery mechanisms; build infrastructure-as-code and automation; improve observability; deliver safe infrastructure migrations; build and evaluate AI-assisted tooling.
Seniority
Principal, hands-on IC with strategy & mentorship

