Infrastructure Engineer/SRE
Core
Designing, building, and advancing core infrastructure to enable engineering teams to execute quickly, productively, and securely for an AI-driven customer experience platform.
Role type
Senior Infrastructure Engineer / Site Reliability Engineer (SRE)
Builds
Developer toolchains, multi-cloud Kubernetes clusters, ML training/deployment infrastructure, and automated operations pipelines.
Domain
Cloud Infrastructure, DevOps, Kubernetes, Machine Learning Infrastructure
Deliverable
infrastructure
Required skills
Golang, Python, Kubernetes, Infrastructure as Code (Terraform/CloudFormation), CI/CD (GitHub Actions), GitOps (Flux/Argo), Cloud Security, PostgreSQL
Preferred skills
GPU-enabled clusters, Google Cloud Platform, Azure, Helm, Kustomize, cert-manager, external-dns
Technologies
Kubernetes, Terraform, CloudFormation, GitHub Actions, Flux, Argo, Helm, Kustomize, PostgreSQL, AWS, Google Cloud, Azure
Responsibilities
Partner with engineers to build dev tools that empower developer workflows and deployment infrastructure; Ensure reliability of multi-cloud Kubernetes clusters and pipelines; Implement metrics, logging, analytics, and alerting for performance and security; Develop Infrastructure-as-code deployment tooling and supporting services; Automate operations and engineering workflows; Build machine learning infrastructure enabling AI teams to train, test, and deploy on large-scale datasets.
Seniority
Senior, hands-on IC
