Cloud Infrastructure Engineer
Core
Design and deploy secure, fault-tolerant cloud infrastructure to support distributed AI workloads, GPU orchestration, and high-availability ML systems for enterprise customers.
Role type
Founding Cloud Infrastructure Engineer (early-stage startup)
Builds
Cloud infrastructure, distributed compute systems, GPU orchestration layers, and observability tooling for AI/ML platforms
Domain
Enterprise AI, Cloud Infrastructure, Distributed Systems
Deliverable
infrastructure
Required skills
Kubernetes, Terraform, distributed systems (Ray, Dask, Spark), Kafka/MQ systems, Python/Go programming, production-grade infrastructure operations
Preferred skills
GPU scheduling/optimization, startup/build-from-scratch experience, supporting ML/AI workloads in production
Technologies
AWS, Azure, GCP, Kubernetes, Terraform, Ray, Kafka, Python, Go
Responsibilities
Architect and scale cloud infrastructure across major providers; Build and optimize compute systems for distributed frameworks; Manage GPU resources and scheduling for ML inference and training; Implement monitoring, security, and best practices for high-availability systems; Collaborate with ML researchers and data scientists to accelerate development
Seniority
Early-stage founding hire, 3–10+ years experience