Member of Technical Staff - Platform Engineering
Core
Building and scaling cloud infrastructure layers to support AI workloads, focusing on reliability, performance, and operational processes for production systems.
Role type
Senior Platform Engineering IC (Infrastructure/Reliability)
Builds
Cloud infrastructure services for AI inference, fine-tuning, and production sandboxes
Domain
Cloud Infrastructure / AI Systems
Deliverable
infrastructure
Required skills
High-quality production code, on-call operations, cloud architecture, auto scaling, fleet management, capacity planning, database operations, monitoring, CI/CD, Kubernetes cluster management
Preferred skills
Systems safety research (STAMP), control theory
Technologies
Kubernetes, Postgres, Redis, AWS
Responsibilities
Identify architectural changes to improve reliability and performance, foster a culture of reliability, define and implement operational processes, operate systems like Kubernetes and databases, participate in on-call rotations and incident response
Seniority
Senior, hands-on IC
