CareerPlanGet AI match score →

Senior Platform Telemetry Engineer

2 Locations💼 Full-time💰 $152,000–$152,000🗓 2026-07-09 → 2026-07-31

Core

Design rack-level solutions for next-generation scaling AI supercomputing platforms and drive fleet management solutions for scaling AI infrastructure using GPUs and Grace.

Role type

Senior Platform Telemetry Engineer

Builds

Next-generation scaling AI supercomputing platforms and fleet management solutions for AI infrastructure.

Domain

AI computing, HPC, GPU infrastructure

Deliverable

production ML models | product features | infrastructure

Required skills

C/C++, Python, time series databases (Influxdb, Prometheus), REST APIs, firmware architecture, system resource analysis, scalability solutions, SCM (Git, Perforce), Jira

Preferred skills

Redfish, notification systems (PagerDuty), x86/ARM system architecture, Confidential Compute, ML and multi-variable optimization

Technologies

NVIDIA GH200, Grace, Influxdb, Prometheus, Grafana, Redfish, PagerDuty, Git, Perforce, Jira

Responsibilities

Design rack-level solutions for scaling AI supercomputing platforms, drive fleet management solutions for AI infrastructure, define architecture for fleet health monitoring and fault-remediation, write architecture specs and design documents, conduct code reviews, ensure product testing and QA, manage product life cycles as product owner, articulate requirements and execution plans.

Seniority

Senior, hands-on IC

Sourced via workday · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Workday ↗