Senior Platform Engineer - Core Infrastructure
Core
Architect, deploy, and operate Kubernetes clusters across bare-metal datacenters to support AI cloud infrastructure for researchers, enterprises, and hyperscalers.
Role type
Senior Platform Engineer (Core Infrastructure)
Builds
Lambda's cloud APIs, systems, and internal tooling for system deployment, management, and maintenance.
Domain
Cloud Infrastructure / Kubernetes / AI Compute
Deliverable
production ML models | infrastructure
Required skills
Kubernetes internals, GitOps, Linux system operations, Infrastructure-as-Code (Terraform/Pulumi), Networking, Observability stacks, Go/Python programming, Security practices
Preferred skills
Multi-cluster/multi-cloud environments, GPU scheduling, Workflow orchestration, Cost optimization, CNCF contributions
Technologies
Kubernetes, Helm, Kustomize, Terraform, Pulumi, Prometheus, Grafana, OpenTelemetry, Go, Python
Responsibilities
Architect and operate Kubernetes clusters; Build automation for cluster lifecycle management; Own reliability, performance, and security of production workloads; Implement observability and alerting; Partner on scalable cloud-native services; Set standards for resource management and networking; Lead incident response and post-mortems; Mentor junior engineers
Seniority
Senior, hands-on IC