Staff Cloud Engineer
Core
Design, architect, and implement scalable, secure, and highly available cloud infrastructure on AWS across multi-account, multi-region environments for financial-services scale.
Role type
Staff Cloud Engineer (AWS & Kubernetes)
Builds
Production EKS clusters, multi-tenant distributed systems, and reusable Infrastructure as Code modules.
Domain
Cloud Infrastructure, Kubernetes, Financial Services
Deliverable
production ML models | product features | infrastructure
Required skills
AWS services (EC2, VPC, IAM, EKS, S3, CloudWatch, API Gateway, Route 53), Terraform, AWS CloudFormation, Docker, Kubernetes architecture, container lifecycle management, Prometheus, Grafana, Dynatrace, OpenSearch, ELK/Loki, Linux/Unix systems administration, Bash, Python, cloud security best practices (IAM, RBAC, secrets management, network security), networking fundamentals (VPCs, subnets, load balancing, DNS, Kubernetes ingress controllers), distributed systems troubleshooting.
Preferred skills
AWS Solutions Architect Professional or DevOps Engineer Professional certification, CKA or CKAD certification, Helm, GitOps tools (ArgoCD, Flux), Rancher, microservices architecture, AWS API Gateway and Lambda Authorizers, cost optimization (Graviton, Spot, Savings Plans), CIAM/identity federation (OIDC, OAuth2, SAML, Auth0), AI/ML infrastructure.
Technologies
AWS, Terraform, AWS CloudFormation, Docker, Kubernetes, EKS, Prometheus, Grafana, Dynatrace, OpenSearch, ELK, Loki, Bash, Python, ArgoCD, Flux, Helm, Rancher, Auth0, Graviton, Spot, Savings Plans, OIDC, OAuth2, SAML.
Responsibilities
Design and implement scalable cloud infrastructure on AWS; define and enforce cloud architecture standards and governance policies; build and maintain IaC using Terraform and AWS CloudFormation; optimize cloud environments for cost, performance, and reliability; design, deploy, and manage production EKS clusters; plan and execute cluster upgrades and lifecycle management; build and maintain internal Helm chart libraries and GitOps-driven configurations; implement zero-trust network principles and IAM least-privilege; drive SRE practices including SLO definition; lead incident response and postmortem analysis; build chaos engineering exercises and disaster recovery testing; evaluate new AWS services and open-source tooling.
Seniority
Staff, hands-on IC with strategic influence