Staff DevOps Engineer
Core
Lead the design and evolution of the platform engineering practice with deep technical ownership of Kubernetes infrastructure to support engineering and AI/ML workloads.
Role type
Staff DevOps Engineer (Individual Contributor)
Builds
Multi-cluster Kubernetes platform, internal developer platform, service mesh, and GPU/ML workload scheduling infrastructure.
Domain
Cloud Infrastructure / Platform Engineering / AI/ML Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Kubernetes architecture and multi-tenancy, GitOps (Argo CD/Flux), Infrastructure-as-Code (Terraform/Pulumi), Service mesh (Istio/Linkerd/Cilium), GPU scheduling and resource management, Cloud infrastructure (AWS/GCP), Observability (Prometheus/Grafana/Datadog/OpenTelemetry), Capacity planning and cost optimization.
Preferred skills
Kubernetes policy-as-code (OPA/Gatekeeper/Kyverno), Multi-cloud experience (Azure), AI/ML infrastructure literacy.
Technologies
Kubernetes, Argo CD, Flux, Terraform, Pulumi, Istio, Linkerd, Cilium, Nginx, Kafka, Redis, Prometheus, Grafana, Datadog, OpenTelemetry, GKE, EKS.
Responsibilities
Own the architecture and roadmap for the multi-cluster Kubernetes platform; Establish GitOps-based deployment workflows; Design and evolve an internal developer platform; Architect service mesh and networking strategies; Define platform standards for GPU and ML workload scheduling; Partner with AI/ML teams on infrastructure requirements; Drive capacity planning and cost optimization; Set technical direction and review architecture for platform-impacting changes; Define and own platform reliability through SLOs and incident-response leadership; Mentor experienced engineers.
Seniority
Staff, Individual Contributor with technical leadership
