Senior Applied Scientist
Core
Building foundational AI technologies for agentic systems, specifically focusing on model post-training and reward modeling to enable self-evolving agents in home shopping.
Role type
Senior Applied Scientist (LLM Post-Training & Reward Modeling)
Builds
Self-evolving agentic systems for home shopping (search, guidance, offer strategy, financing)
Domain
Real Estate / Generative AI / Reinforcement Learning
Deliverable
production ML models
Required skills
LLM post-training (SFT, DPO, RFT/GRPO), Reward model development, Generative AI (transformers, RL, preference learning), Python, PyTorch or TensorFlow
Preferred skills
PhD in CS/ML, Agentic AI evaluation (LLM-as-a-Judge), Published work in RLHF/RLAIF, GPU training platforms (Databricks, Fireworks)
Technologies
PyTorch, TensorFlow, Databricks, Fireworks
Responsibilities
Own LLM post-training pipelines (SFT, DPO, RFT/GRPO) on GPU infrastructure; Build and train multi-category reward models (PRMs); Design on-policy and online assessment using LLM-as-a-Judge; Translate offline evaluation rubrics into generalizable reward functions; Provide technical leadership and mentorship to scientists and MLEs
Seniority
Senior, hands-on IC with technical leadership
