Staff Machine Learning Research Engineer, Agent Post-training - Enterprise GenAI
Core
Building next-gen Agent RL training platforms to train state-of-the-art agents for complex enterprise use-cases.
Role type
Staff Machine Learning Research Engineer (Agent Post-training)
Builds
Next-generation AI cybersecurity firewall LLMs and foundation healthtech search models for enterprise clients.
Domain
Generative AI, Reinforcement Learning, Enterprise AI
Deliverable
production ML models
Required skills
LLM training in production, post-training methods (RLHF/RLVR), RL algorithms (PPO/GRPO), multi-agent system design, reward modeling
Preferred skills
Publications in top conferences (NEURIPS, ICLR, ICML)
Technologies
RLHF, RLVR, PPO, GRPO
Responsibilities
Train state-of-the-art models for enterprise deployment, research and integrate cutting-edge algorithms into the training stack, design solutions for complex multi-agent systems learning from process and outcome rewards
Seniority
Staff, hands-on IC with research leadership