ML Infrastructure Engineer
Core
Building and optimizing the reliable, high-performance ML platform that powers recommendations on X.
Role type
ML Infrastructure Engineer
Builds
GPU compute infrastructure, training frameworks, experimentation tools, data pipelines, and large-scale ML systems
Domain
AI/ML infrastructure for social media recommendations
Deliverable
production ML models | infrastructure
Required skills
Python, C++ or Rust, distributed systems, GPU infrastructure, deep learning applications, ML platforms, training infrastructure
Preferred skills
JAX or PyTorch, Linux systems, orchestration tools, job schedulers (Slurm), configuration management (Puppet/Ansible)
Technologies
Python, C++, Rust, JAX, PyTorch, CUDA, Slurm, Puppet, Ansible
Responsibilities
Designing, building, and scaling GPU compute infrastructure; Developing data pipelines and integrating large-scale data, training, and inference systems; Collaborating with ML teams to productionize models; Ensuring scalability, reliability, and efficiency of large-scale machine learning systems; Working across the full stack to solve complex problems independently; Mentoring junior engineers
Seniority
Mid-level, hands-on IC