Forward Deployed Engineer, Compute Operations
Core
Build autonomous software systems for fleet health, automated repair/RMA workflows, hardware qualification, and facility operations to manage gigawatt-scale AI compute infrastructure.
Role type
Senior IC software engineer (compute operations & infrastructure)
Builds
Real-time telemetry systems, automated incident response, hardware validation pipelines, and asset management tools for data centers.
Domain
AI infrastructure / Data Center Operations / Hardware-Software Integration
Required skills
Go, Python, TypeScript, Kubernetes, bare metal management, LLM API integration, agentic frameworks, autonomous coding tools, on-call rotation experience, system design, operational workflow automation.
Preferred skills
GPU fleet management, BMC/Redfish/IPMI tooling, CMMS/DCIM systems, BMS/EPMS/SCADA, Prometheus, Grafana.
Responsibilities
Build fleet health systems with real-time telemetry and automated incident correlation; automate repair and RMA flows from failure detection to return-to-service; ship hardware qualification as repeatable software workflows; run facility maintenance systems and migrate legacy datacenter inventory; convert runbooks and SOPs into structured, auditable data and dashboards. (via careerplan.io/jobs/3dacc390-e34f-4644-8138-3dbd32d68682-forward-deployed-engineer-compute-operations-at-fluidstack)
Seniority
Senior, hands-on IC
