论文

生成前规范:对金钱、时间、幂等性和访问任务中 LLM 生成代码的规范框架进行预注册、五模型配对评估

Specification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks

上下文与知识上下文工程

摘要

Code generated by large language models passes security checks at a rate that has barely moved in four years. In regulated backends, the defect classes that matter most are money arithmetic, time handling, retry safety, and access control.团队用指令文件来回答,但迄今为止最大规模的指令文件对照研究发现没有任何好处。本文测试了一个更狭隘的想法:当提示携带规范时,生成的代码会得到改进,这是一个固定的前导码,说明结果必须是正确的。我们预先注册了假设、反驳器、分析代码和一次性生成规则,然后通过来自五个供应商谱系的五个前沿模型运行来自金融、医疗保健和保险实践的 50 个实际后端任务,每个任务两次:裸露,前面是一个 267 字的填充规范框架。 Nine deterministic AST-based checkers scored the outputs. The Bandit security scanner, which knows nothing of the frame, scored them independently.该框架减少了所有五个模型的缺陷(每个任务平均减少 0.16 至 0.70 个结果,每个 Holm 调整符号检验显着,每个引导置信区间不包括零)。在手臂不同的地方,框架臂赢了 100 次中的 95 次。它从未使任何领域的任何模型变得更糟。 Bandit found 53 medium-or-high issues in the bare arm and 11 in the frame arm, in the same direction for every model. The effect was largest where a model's unprompted defaults were weakest: the frame supplies the discipline a model lacks.所有 500 个输出、提示、检查器、评分代码和预注册均通过 DOI 发布,因此任何团队都可以重新得出结果,而无需信任作者。