使用本地编码代理推广操作技能
Generalizing Manipulation Skills with a Local Coding Agent
摘要
如今,开放权重语言模型的进步使系统能够在单个工作站上运行的同时编写、执行和调试代码。 Most language-driven robots give the model a fixed action interface or a trained policy. Generalizing to a new task therefore means more engineering effort or more data collection, both time-consuming.我们研究本地开放权重视觉语言模型是否可以控制机器人并一次性推广到任务的新变体,而无需新的人类编程或训练。 We let a local open-weight VLM, Qwen3.8-27B, drive a UR3e robotic arm from a coding-agent harness. It writes and runs its own code above a service that implements kinematics, safety limits and classic computer vision techniques. We investigate if this system is capable of generalizing to unseen tasks.具体来说,我们在由儿童玩具构建的九个任务上对其进行了测试,这些任务旨在探索各种对象特征的泛化能力:颜色、大小、形状和这些对象的任务变化。每个任务进行 5 次试验,我们在 45 次试验中观察到 30 次的泛化情况,持续时间从 3.4 到 67.5 分钟不等,具体取决于任务复杂性。 We further test if there is a speedup when an agent is asked to redo the task after successful completion. This resulted in a 50% reduction in duration, indicating that there is self-improvement over time. Finally, we expose the limitations of a local coding agent.我们相信,解决这些限制并结合随着时间的推移对自我改进的进一步研究,可以直接通往本地编码代理的实际部署。