Principal Member of Technical Staff
Core
Architect, implement, and operate highly available Kubernetes clusters supporting thousands of concurrent persistent resources and AI agent lifecycles.
Role type
Principal Member of Technical Staff (Platform/Infrastructure)
Builds
Highly available Kubernetes clusters, custom resource definitions, operators, and reproducible infrastructure-as-code environments.
Domain
Cloud Infrastructure / Kubernetes / AI Research Platforms
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals (API server, etcd, scheduler, controller manager, kubelet), CRDs, Kubernetes operators, Terraform/Pulumi/Crossplane, container networking (CNI, service mesh, CSI), AWS EKS/GCP GKE/Azure AKS, systems/backend programming languages, GitOps workflows, observability and incident response.
Preferred skills
Data-intensive platforms, scientific computing, ML/AI infrastructure, startup scaling experience, significant architectural ownership.
Responsibilities
Design and develop custom resource definitions and Kubernetes operators for AI agent lifecycles; define cluster scaling, node pool management, and autoscaling strategies; build and maintain infrastructure-as-code environments; establish observability and monitoring practices; troubleshoot complex distributed infrastructure issues; collaborate with backend and ML teams to translate workload requirements into infrastructure patterns.
Seniority
Principal, hands-on IC with strategic ownership
