Rewrite
## About the Role
We're looking for a Senior Infrastructure & DevOps Engineer who brings deep hands-on experience with AWS, Kubernetes, automation, and large-scale distributed data systems.
You should be someone who thrives in fast-paced environments, enjoys ownership, communicates clearly, and consistently produces well-structured documentation.
This is a highly impactful role. You will be a key contributor shaping the stability, scalability, and velocity of Carbon Arc's data platform, the lakehouse, query engines, graph, and the applications built on top of them.
## Responsibilities
### Core Infrastructure & Cloud Operations
- Own, improve, and maintain AWS workloads across EKS, EC2, S3, RDS (Postgres), ElastiCache/Redis, Lambda, IAM, VPC, and ALB/NLB ingress.
- Collaborate on architecture decisions for reliability, cost optimization, multi-account strategy, and security posture.
- Manage and optimize large-scale object storage (S3) backing an Apache Iceberg lakehouse, and keep data-intensive workloads performant and cost-efficient.
- Own infrastructure-as-code in Terraform / Terragrunt, including AWS resources, IAM boundaries, and GitHub org configuration.
### Kubernetes Operations
- Operate and harden EKS clusters hosting StarRocks, Apache Airflow, Neo4j, web applications and APIs, plus StatefulSets and persistent volumes.
- Develop reliable Kubernetes deployment patterns using Helm charts and GitOps via ArgoCD (our developers-argocd repository).
- Manage secret delivery with External Secrets Operator backed by AWS Secrets Manager, and ingress via the AWS Load Balancer Controller.
- Maintain multi-namespace, multi-environment (dev / stage / prod) conventions and feature-branch preview environments.
### Platform Reliability & Automation
- Build automated systems for deployments, scaling and autoscaling, monitoring and alerting, disaster recovery, and infrastructure-as-code.
- Strengthen our CI/CD on GitHub Actions: container builds to ECR, GitOps-driven deploys, and preview environments.
- Improve observability via Datadog (with room to expand into tooling such as Prometheus and OpenTelemetry).
- Support the data platform's compute layer: Apache Spark, Domino Data Lab job runtimes, dbt, and Airflow DAG orchestration.
### Documentation & Communication
- Produce clear, thorough documentation: runbooks, architecture diagrams, and internal guides.
- Communicate proactively with engineering, product, and leadership.
- Participate in incident response, post-mortems, and continuous improvement.
## Required Experience
### Cloud & AWS
- 5+ years hands-on with AWS.
- Deep understanding of:
- IAM and identity boundaries
- EKS operations
- VPC networking, peering, and Transit Gateway
- S3 performance and lifecycle optimization
- RDS (Postgres)
### Kubernetes
- Advanced operational experience with:
- Deployments, StatefulSets, DaemonSets
- Persistent storage and stateful workloads
- Ingress, ALB/NLB, and service patterns
- Multi-namespace and multi-environment best practices
- Debugging production containers and pods
- Hands-on experience with Helm and a GitOps tool (ArgoCD or Flux).
### Infrastructure-as-Code & CI/CD
- Production experience with Terraform (Terragrunt a strong plus).
- Building and maintaining CI/CD pipelines (GitHub Actions preferred).
- Container build and registry workflows (Docker, ECR).
## Soft Skills
- Excellent written and verbal communication.
- Organized, structured documentation habits.
- Strong ownership and initiative.
- Calm under pressure.
- Ability to work independently and suggest improvements.
## Nice to Have
- Operating data engines in production: Trino, StarRocks, or Spark.
- Apache Iceberg / lakehouse and catalog (e.g. Polaris) experience.
- Airflow, dbt, or Neo4j operations.
- Okta / SSO integration, ECS, and compliance frameworks (SOC 2, Vanta).
## What Success Looks Like
Within 3 months, you should:
- Take ownership of key AWS and EKS components.
- Produce high-quality documentation.
- Identify reliability improvements.
- Deploy or maintain major internal services (StarRocks, Neo4j, Airflow, web apps, and APIs).
- Improve cost visibility and alerting.
Within 6 to 12 months, you should:
- Lead large-scale infrastructure projects.
- Provide architectural recommendations.
- Harden monitoring, disaster recovery, and cross-VPC communication.
- Accelerate deployment velocity and reliability across the platform.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.