Distributed Software Engineer
Core
Building and operating large-scale AI supercomputer clusters using Wafer-Scale Engine (WSE) chips to enable ultra-high-speed AI training and inference.
Role type
Distributed Systems Software Engineer (Cluster Infrastructure)
Builds
Wafer-Scale Cluster software, orchestration schedulers, monitoring systems, and deployment tools for multi-exaflop supercomputers.
Domain
AI Infrastructure / High-Performance Computing / Distributed Systems
Deliverable
infrastructure
Required skills
distributed systems architecture, Kubernetes ecosystem, GoLang, Python, bash, system design, debugging distributed systems, test automation
Preferred skills
bare-metal configuration, high availability design, cloud and on-premise deployment patterns
Technologies
Kubernetes, Prometheus, Grafana, GoLang, Python, bash
Responsibilities
Automate bare-metal configuration of networking, OS, and application software in large clusters; develop push-button workflows for cluster upgrades, downgrades, and security patching; build orchestration and scheduler systems for resource allocation and job submission; support on-premise and cloud mode deployment and operations; design robust systems for monitoring, detecting, and handling failures including High Availability; develop user-facing and administrator-facing tools for job monitoring, metrics collection, and cluster management.
Seniority
Mid-Senior, hands-on IC