Principal Software Engineer
Core
Design, test, and maintain large-scale research platforms and AI infrastructure (compute clusters) to optimize uptime, performance, and accelerate research pace including agentic development.
Role type
Principal Software Engineer (AI Infrastructure & HPC)
Builds
End-to-end systems, libraries, tools, and AI infrastructure roadmaps for compute clusters and research platforms.
Domain
Artificial Intelligence Infrastructure / High-Performance Computing (HPC)
Deliverable
production ML models | infrastructure
Required skills
Distributed computing, Platform design & delivery, Large-scale ML/AI infrastructure, Agentic workflows, System profiling & benchmarking, GPU programming, Networking (InfiniBand, NVLink), Storage systems, Distributed training parallelisms
Preferred skills
Kubernetes, Docker, Volcano, SLURM, CUDA, NCCL, PyTorch, C, C++, C#, Java, JavaScript, Python
Responsibilities
Design and maintain large-scale research platforms driving compute clusters; Develop end-to-end systems and tools to accelerate research; Gather data to develop AI Infrastructure roadmap; Provide technical leadership and mentorship; Collaborate cross-discipline with engineers, PMs, and science teams; Profile, benchmark, and optimize performance-critical systems.
Seniority
Principal, strategy & mentorship