Research Engineer, Frontier Capabilities
Core
Design, build, and optimize systems for training large language models to run long-horizon scientific discovery tasks, spanning post-training stacks from SFT to asynchronous RL.
Role type
Research Engineer (Frontier Capabilities)
Builds
Post-training infrastructure, agentic harnesses, and scientific benchmarks for autonomous scientific discovery.
Domain
Artificial Intelligence / Scientific Discovery
Deliverable
production ML models | infrastructure | dashboards & analysis
Required skills
Python, distributed ML training frameworks, large-scale model training techniques, cloud or HPC environment, software engineering
Preferred skills
C++/CUDA, open-source ML framework contributions, RL post-training, MoE architectures, large-scale scientific datasets
Technologies
Megatron-LM, TorchTitan, DeepSpeed, Ray
Responsibilities
Profile and optimize GPU utilization for 100B+ parameter training runs; build modular post-training workflows; lead experimentation on reasoning model development; design scientific agentic benchmarks; train models for planning and tool use over extended horizons
Seniority
Individual Contributor (IC), hands-on engineering