Senior Full Stack Software Engineer
Core
Designing and developing a massively distributed scalable platform to identify, diagnose, and remediate non-performant GPU assets for large-scale AI clusters.
Role type
Senior Full Stack Software Engineer (AI Infrastructure)
Builds
Production systems enabling large scalable GPU clusters for AI workloads
Domain
AI computing, GPU infrastructure, distributed systems
Deliverable
production ML models | infrastructure
Required skills
React, TypeScript/JavaScript, Golang, SQL databases, cluster management, incident management, asynchronous workflows, event-driven architecture
Preferred skills
Managing and automating large-scale distributed systems independent of cloud providers, deep understanding of Kubernetes/Slurm, LLM usage and safety
Technologies
React, Web Components, TypeScript, Golang, PostgreSQL, Temporal, Bazel, Kubernetes
Responsibilities
Designing and developing scalable platforms for GPU asset management, ensuring production AI cluster reliability and performance, evaluating system failures, working across the product stack
Seniority
Senior, hands-on IC