Figure 1 : Overview of Dream-RSI . The system operates in a recursive self-improvement loop via three core stages: ① Online Explore , where the current exploration policy guides a coding agent to expand a discovery tree and log historical traces; ② Construct Replay Simulator , where the generated discovery tree is converted into a reusable simulator pool; and ③ Dreaming-based Policy Improvement , where the agent "dreams" up a massive pool of alternative policies in its mind. It then feeds these candidate policies into the replay simulator to simulate executions and derive rapid feedback, continuously refining its strategy (detailed in the Zoom-in box). The updated policy then redeploys for the next round of online exploration.