Staff Software Engineer, MetalDev
Core
Build Go-based distributed services and automation for large-scale GPU data center infrastructure, managing hardware lifecycle, health monitoring, and fleet operations.
Role type
Staff Software Engineer (Infrastructure/Compute Architecture)
Builds
Go-based distributed services, automation workflows, APIs, and observability tooling for GPU server fleets.
Domain
Cloud Infrastructure / Data Center Operations / GPU Compute
Deliverable
production ML models | infrastructure
Required skills
Go, distributed systems, Kubernetes, API design (REST/gRPC), observability (Prometheus/Grafana), incident response, technical leadership, hardware lifecycle management
Preferred skills
Kafka, ClickHouse, CRDB, DMTF/RedFish APIs, GPU server expertise
Technologies
Go, Kubernetes, Prometheus, Grafana, PromQL, gRPC, REST, Kafka, ClickHouse, CRDB
Responsibilities
Design and operate Go services for GPU data center lifecycle management; build automation for bring-up, discovery, and remediation; develop APIs for BMC and firmware state management; improve observability and alerting tooling; translate hardware failure modes into software resilience improvements; provide technical leadership via design/code reviews and mentorship.
Seniority
Staff, hands-on IC with strategic influence