CareerPlanGet AI match score →

AI Research Engineer

💼 Full-time🗓 2026-06-24

Core

Design and develop next-generation agentic AI systems, defining how agents reason, remember, and improve over time.

Role type

Senior to Principal-level AI Research Engineer (Agentic Systems)

Builds

Scalable agentic AI systems serving Dropzone AI's product capabilities

Domain

Applied AI, Agentic Systems, LLMs

Deliverable

production ML models

Required skills

Agent architecture design, Harness and memory engineering, Robust evaluation and benchmarking, Multi-step reasoning agent implementation, Multi-agent coordination frameworks, Memory subsystem architecture, Automated evaluation pipeline development, Research-to-production translation

Preferred skills

Context/harness engineering mindset, Ability to test latest research in real-world deployment, Converting non-deterministic LLM outputs to consistent outcomes, Replicating expert human intuitions, Ownership mindset, Driving ambiguous problems to clarity

Technologies

LLMs, Agentic frameworks, Memory systems, Evaluation frameworks

Responsibilities

Design and implement advanced multi-step reasoning agents with tool use, planning, and self-improvement loops; Architect short-term and long-term memory subsystems including episodic, semantic, and retrieval-based mechanisms; Define and implement evaluation frameworks for agent performance including task success and reasoning quality; Translate latest community research ideas into production-grade systems and run experiments to iterate quickly.

Rewrite
## About the role We are seeking a Senior to Principal-level AI Research Engineer to lead the design and development of next-generation agentic AI systems. This role sits at the intersection of research and production, with a strong emphasis on: - Agent architecture design - Harness and memory engineering - Robust evaluation and benchmarking of model and agent performance You will work closely with product and engineering teams to translate cutting-edge research into scalable, real-world systems. In this role, you will directly shape the core intelligence layer of Dropzone AI. Your work will define how our agents reason, remember, and improve over time, influencing both our product capabilities and the broader direction of applied AI systems. ## What we're looking for - Someone who thinks in context/harness engineering, not just models - A learner who can follow latest research and test them in real-world deployment - Deep curiosity about how to convert non deterministic outputs from LLMs to consistent reliable outcomes and replicate expert human intuitions - Strong ownership mindset and ability to drive ambiguous problems to clarity ## What you'll do ### Agentic Architecture - Design and implement advanced multi-step reasoning agents (tool use, planning, reflection, self-improvement loops) - Develop frameworks for multi-agent coordination and task decomposition - Improve reliability, latency, and cost efficiency of agent execution ### Memory Systems - Architect short-term and long-term memory subsystems (episodic, semantic, retrieval-based, hybrid) - Build mechanisms for context compression, retrieval, and grounding - Explore novel approaches to continual learning and state persistence ### Evaluation & Reliability - Define and implement evaluation frameworks for agent performance (task success, reasoning quality, robustness) - Build automated eval pipelines (synthetic data, adversarial testing, regression testing) - Establish metrics and benchmarks for agent reliability in production ### Research → Production - Translate latest community research ideas into production-grade systems - Run experiments, analyze results, and iterate quickly - Contribute to internal
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗