Software Engineer, Fleet Automation
Core
Design and build automation platforms, internal services, and APIs to provision, configure, and manage large-scale GPU and CPU compute nodes at peak efficiency.
Role type
Senior IC backend software engineer (infrastructure automation)
Builds
Production fleet automation services, orchestration layers, and observability dashboards for HPC infrastructure
Domain
High-performance computing (HPC) and cloud infrastructure
Deliverable
production ML models | product features | infrastructure
Required skills
Go, C#, TypeScript, Linux systems administration, relational and NoSQL database design, CI/CD pipeline engineering, observability stack implementation, incident response, backend service architecture
Preferred skills
GPU compute infrastructure knowledge, NVIDIA tooling (DCGM, nvidia-smi), event-driven architectures, Kafka messaging platforms
Technologies
Go, C#, TypeScript, Prometheus, Grafana, Alertmanager, ELK, Kafka, Ubuntu, RHEL, NVIDIA Container Toolkit
Responsibilities
Design and maintain fleet automation services for lifecycle management of compute nodes; Develop APIs enabling hardware deployment and decommissioning; Build backend services with focus on reliability and maintainability; Design data models for automation workflows; Maintain CI/CD pipelines for hardware validation; Instrument systems for real-time fleet health visibility; Participate in on-call rotations and own incident response
Seniority
Senior, hands-on IC