CareerPlanGet AI match score →

Sr. Infrastructure and DevOps Engineer

Onsite or remote • New York City+1🌐 Remote💼 Full-time🗓 2026-06-25

Core

Design, build, and maintain the stability, scalability, and security of a large-scale distributed data platform (lakehouse, query engines, graph) and supporting applications on AWS.

Role type

Senior Infrastructure & DevOps Engineer (Cloud & Kubernetes)

Builds

AWS workloads, EKS clusters, data lakehouse infrastructure, CI/CD pipelines, and observability systems.

Domain

Cloud Infrastructure, Kubernetes, Data Engineering

Deliverable

production ML models | infrastructure

Required skills

AWS (EKS, VPC, IAM, S3, RDS), Kubernetes (StatefulSets, Helm, GitOps), Terraform, CI/CD (GitHub Actions), Observability (Datadog), Data Platform (Apache Iceberg, Spark, Airflow)

Preferred skills

StarRocks, Trino, Apache Iceberg, Polaris, Neo4j, Okta/SO2 compliance

Technologies

AWS, EKS, Terraform, Terragrunt, Helm, ArgoCD, GitHub Actions, Datadog, Apache Iceberg, StarRocks, Apache Spark, Airflow, Neo4j

Responsibilities

Own and optimize AWS workloads and EKS clusters; Manage large-scale object storage and data-intensive workloads; Build and maintain IaC and CI/CD pipelines; Operate data engines and orchestration tools; Produce documentation and lead incident response.

Seniority

Senior, hands-on IC

Rewrite
## About the Role We're looking for a Senior Infrastructure & DevOps Engineer who brings deep hands-on experience with AWS, Kubernetes, automation, and large-scale distributed data systems. You should be someone who thrives in fast-paced environments, enjoys ownership, communicates clearly, and consistently produces well-structured documentation. This is a highly impactful role. You will be a key contributor shaping the stability, scalability, and velocity of Carbon Arc's data platform, the lakehouse, query engines, graph, and the applications built on top of them. ## Responsibilities ### Core Infrastructure & Cloud Operations - Own, improve, and maintain AWS workloads across EKS, EC2, S3, RDS (Postgres), ElastiCache/Redis, Lambda, IAM, VPC, and ALB/NLB ingress. - Collaborate on architecture decisions for reliability, cost optimization, multi-account strategy, and security posture. - Manage and optimize large-scale object storage (S3) backing an Apache Iceberg lakehouse, and keep data-intensive workloads performant and cost-efficient. - Own infrastructure-as-code in Terraform / Terragrunt, including AWS resources, IAM boundaries, and GitHub org configuration. ### Kubernetes Operations - Operate and harden EKS clusters hosting StarRocks, Apache Airflow, Neo4j, web applications and APIs, plus StatefulSets and persistent volumes. - Develop reliable Kubernetes deployment patterns using Helm charts and GitOps via ArgoCD (our developers-argocd repository). - Manage secret delivery with External Secrets Operator backed by AWS Secrets Manager, and ingress via the AWS Load Balancer Controller. - Maintain multi-namespace, multi-environment (dev / stage / prod) conventions and feature-branch preview environments. ### Platform Reliability & Automation - Build automated systems for deployments, scaling and autoscaling, monitoring and alerting, disaster recovery, and infrastructure-as-code. - Strengthen our CI/CD on GitHub Actions: container builds to ECR, GitOps-driven deploys, and preview environments. - Improve observability via Datadog (with room to expand into tooling such as Prometheus and OpenTelemetry). - Support the data platform's compute layer: Apache Spark, Domino Data Lab job runtimes, dbt, and Airflow DAG orchestration. ### Documentation & Communication - Produce clear, thorough documentation: runbooks, architecture diagrams, and internal guides. - Communicate proactively with engineering, product, and leadership. - Participate in incident response, post-mortems, and continuous improvement. ## Required Experience ### Cloud & AWS - 5+ years hands-on with AWS. - Deep understanding of: - IAM and identity boundaries - EKS operations - VPC networking, peering, and Transit Gateway - S3 performance and lifecycle optimization - RDS (Postgres) ### Kubernetes - Advanced operational experience with: - Deployments, StatefulSets, DaemonSets - Persistent storage and stateful workloads - Ingress, ALB/NLB, and service patterns - Multi-namespace and multi-environment best practices - Debugging production containers and pods - Hands-on experience with Helm and a GitOps tool (ArgoCD or Flux). ### Infrastructure-as-Code & CI/CD - Production experience with Terraform (Terragrunt a strong plus). - Building and maintaining CI/CD pipelines (GitHub Actions preferred). - Container build and registry workflows (Docker, ECR). ## Soft Skills - Excellent written and verbal communication. - Organized, structured documentation habits. - Strong ownership and initiative. - Calm under pressure. - Ability to work independently and suggest improvements. ## Nice to Have - Operating data engines in production: Trino, StarRocks, or Spark. - Apache Iceberg / lakehouse and catalog (e.g. Polaris) experience. - Airflow, dbt, or Neo4j operations. - Okta / SSO integration, ECS, and compliance frameworks (SOC 2, Vanta). ## What Success Looks Like Within 3 months, you should: - Take ownership of key AWS and EKS components. - Produce high-quality documentation. - Identify reliability improvements. - Deploy or maintain major internal services (StarRocks, Neo4j, Airflow, web apps, and APIs). - Improve cost visibility and alerting. Within 6 to 12 months, you should: - Lead large-scale infrastructure projects. - Provide architectural recommendations. - Harden monitoring, disaster recovery, and cross-VPC communication. - Accelerate deployment velocity and reliability across the platform.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗