论文

Dream-RSI:通过历史探索回放改进Agent搜索策略

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

智能体系统Agent 规划Agent Harness智能体自我改进递归自我改进(RSI)

摘要

Dream-RSI把探索策略作为可修改的代码:决定继续哪些分支、并行开展多少尝试,以及何时停止。负责生成候选方案的编码Agent保持不变,更新的是组织探索的方法。 系统把已执行的探索过程整理成发现树,保存分支和实际结果。新的策略可以在这些记录上回放,比较发现质量与探索开销,减少为评估策略而重复调用编码Agent和执行评测。选出的策略投入下一轮真实探索,新产生的记录继续加入回放环境。 首版在算法工程、数学优化和GPU内核优化中比较固定探索策略与持续改进策略,观察到部分任务达到相近质量所需的尝试更少,另一些任务在相近预算下得到更好的结果。回放只读取已有分支的已知结果,新路径仍由真实探索产生。

Dream-RSI的探索与历史回放改进流程
Figure 1 : Overview of Dream-RSI . The system operates in a recursive self-improvement loop via three core stages: ① Online Explore , where the current exploration policy guides a coding agent to expand a discovery tree and log historical traces; ② Construct Replay Simulator , where the generated discovery tree is converted into a reusable simulator pool; and ③ Dreaming-based Policy Improvement , where the agent "dreams" up a massive pool of alternative policies in its mind. It then feeds these candidate policies into the replay simulator to simulate executions and derive rapid feedback, continuously refining its strategy (detailed in the Zoom-in box). The updated policy then redeploys for the next round of online exploration.