Member of Engineering (Scalability)
Core
Building distributed training and inference infrastructure for Large Language Models (LLMs) to enable source code generation.
Role type
Senior IC distributed systems engineer (LLM infrastructure)
Builds
High-reliability distributed training and inference systems for foundational LLMs
Domain
AI infrastructure / Large Language Models / High-performance computing
Deliverable
production ML models
Required skills
Linux kernel debugging, distributed systems, fault tolerance, NCCL, PyTorch, C/C++, CUDA API, algorithmic design
Preferred skills
GPU architecture knowledge, checkpointing optimization, observability, K8s stack
Technologies
PyTorch, CUDA, NCCL, Linux kernel, C/C++, Python, Cython, K8s
Responsibilities
Troubleshoot hardware problems during large-scale training, minimize GPU idle time during faults, design tools for training recovery, improve checkpointing performance and reliability, write high-performance code in Python/C/C++/CUDA
Seniority
Senior, hands-on IC