Software Engineer, Platform
Core
Build software and automation to manage a large fleet of GPU servers, ensuring high availability and performance for AI workloads.
Role type
Senior IC Platform Engineer (GPU Infrastructure)
Builds
Automated server provisioning, health monitoring, error detection, and recovery tooling for GPU fleets.
Domain
Cloud Infrastructure / AI Hardware Operations
Deliverable
infrastructure
Required skills
Python, Linux systems administration, GPU infrastructure management, storage systems, infrastructure-as-code, network diagnostics
Preferred skills
NVIDIA GPU stack optimization, AMD GPU experience, bare-metal provisioning, compliance frameworks
Technologies
Python, Ansible, Terraform, NVMe, NFS, Lustre, GPFS, NVIDIA drivers, CUDA, DCGM, InfiniBand, RoCEv2, BGP, tcpdump
Responsibilities
Build Python fleet tracking systems for server lifecycle management; develop automated provisioning and recovery tooling; create metrics and dashboards for hardware health; tune Linux systems for AI workloads; implement OS-level security hardening; manage distributed storage systems.
Seniority
Senior, hands-on IC