Research Engineer, Discovery
Core
Design and implement large-scale infrastructure systems to support AI scientist training, evaluation, and deployment across distributed environments.
Role type
Senior IC research engineer (ML infrastructure)
Builds
Large-scale data pipelines, VM/sandboxing/container architectures, and evaluation frameworks for scientific AGI
Domain
AI safety and alignment research infrastructure
Deliverable
infrastructure
Required skills
large-scale distributed systems, performance optimization, containerization (Docker, Kubernetes), data pipeline engineering, complex infrastructure debugging, cross-functional collaboration
Preferred skills
language model training infrastructure, distributed ML frameworks (PyTorch, JAX), GPU/TPU architecture knowledge, cloud platforms (AWS, GCP), workflow orchestration, reinforcement learning
Technologies
PyTorch, JAX, Docker, Kubernetes, AWS, GCP, Beam, Spark, Dask
Responsibilities
Design and implement large-scale infrastructure systems for AI scientist training and deployment; Identify and resolve infrastructure bottlenecks; Develop robust evaluation frameworks for scientific AGI progress; Build scalable VM/sandboxing/container architectures; Translate experimental requirements into production-ready infrastructure; Optimize large-scale training and inference pipelines for reinforcement learning
Seniority
Senior, hands-on IC