CareerPlanSign in

Software Engineer, Compute Operations

New York, NY💼 Full-time🗓 2026-10-02 → 2026-10-07

Core

Building fleet health systems, automated repair workflows, and hardware qualification software for large-scale AI compute infrastructure.

Role type

Senior IC software engineer (compute operations)

Builds

Real-time telemetry systems, automated RMA flows, hardware validation workflows, and facility maintenance systems for GPU fleets.

Domain

AI infrastructure / Data center operations / Hardware management

Required skills

Go, Python, TypeScript, Kubernetes, bare metal management, LLM API integration, agentic frameworks, autonomous agent development, on-call rotation experience, production engineering, SRE practices, hardware qualification, burn-in frameworks, BMC/Redfish/IPMI tooling, CMMS/DCIM systems, BMS/EPMS/SCADA, Prometheus, Grafana.

Preferred skills

Production engineering on large GPU fleets, hardware qualification frameworks, BMC/Redfish/IPMI tooling, CMMS/DCIM systems, BMS/EPMS/SCADA, Prometheus, Grafana.

Technologies

Kubernetes, Go, Python, TypeScript, OpenAI, Anthropic, MCP servers, Claude Code, Cursor, Prometheus, Grafana, BMC, Redfish, IPMI, CMMS, DCIM, BMS, EPMS, SCADA. (via careerplan.io/jobs/d724aa9e-db4f-443e-8f86-9d8ac13fd625-software-engineer-compute-operations-at-fluidstack)

Responsibilities

Build real-time telemetry and healthcheck systems for machines across Kubernetes and bare metal; automate repair and RMA processes from failure detection to return to service; ship hardware qualification as software including burn-in and performance baselining; run facility maintenance systems for lockout tagout and work orders; convert runbooks and SOPs into structured, auditable procedures.

Seniority

Senior, hands-on IC