论文

SWE-Skills-Bench:用真实软件开发任务检验Agent Skills的增益

SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?

智能体系统模型评测应用与实践Agent任务评测Agent Skills基准与评测资源编程

摘要

Agent Skills把操作方法组织为可加载的知识包,但是否能改善真实软件开发任务,需要与不加载技能的基线比较。SWE-Skills-Bench将49项公开技能与固定提交版本的GitHub仓库、明确验收条件的需求文档配对,在六类开发场景中构建约565项任务,并将验收条件转换为可执行测试。 在论文评测中,39项技能没有提高任务通过率,平均增益约1.2%;七项专业技能带来明确收益,三项因指导与项目版本不匹配而降低表现。部分技能显著增加Token消耗,却没有改善通过率。结果说明,技能的专业针对性、适用版本和当前项目上下文,比是否加载技能本身更能解释效果。该论文主分类为cs.SE,本次作为Agent Skills研究的重要负向证据定向补入。

技能选择、程序知识加载、工具执行与结果验收流程
Figure 1 : Illustration of how agent skills are used in a software engineering workflow. Given a natural-language requirement, the LLM-based agent selects the most relevant skill from its skill library, including skills such as writing code, running tests, debugging, creating pull requests, and deploying, and injects it into the context window. The agent then executes a series of SWE actions to produce the final software artifacts (such as code) that fulfill the requirement.