Infrastructure Engineer/SRE
Core
Designing, building, and advancing core infrastructure to enable engineering teams to execute quickly, productively, and securely, including ML infrastructure for AI teams.
Role type
Senior Infrastructure Engineer / Site Reliability Engineer (SRE)
Builds
Developer toolchains, multi-cloud Kubernetes clusters, CI/CD pipelines, and ML training/deployment infrastructure.
Domain
Cloud Infrastructure, DevOps, Machine Learning Operations
Deliverable
infrastructure
Required skills
Golang, Python, Kubernetes, Infrastructure as Code (Terraform/CloudFormation), CI/CD (GitHub Actions), GitOps (Flux/Argo), PostgreSQL, container security
Preferred skills
GPU-enabled clusters, multi-cloud (Google Cloud, Azure)
Technologies
Kubernetes, Helm, Kustomize, Terraform, CloudFormation, AWS (IAM, S3, EC2, EKS), GitHub Actions, Flux, Argo, PostgreSQL
Responsibilities
Partner with engineers to build dev tools and deployment infrastructure; ensure reliability of multi-cloud Kubernetes clusters and pipelines; implement metrics, logging, analytics, and alerting; develop IaC deployment tooling; automate operations and engineering workflows; build ML infrastructure for large-scale datasets.
Seniority
Senior, hands-on IC
