论文

Viva La Vida:验证和多智能体证明搜索中的累积失败

Viva La Vida: Verification and Accumulation Failures in Multi-Agent Proof Search

智能体系统Agent任务评测

摘要

当Agentic证明器处理一个开放问题时,没有可以依靠的证明助手:它的验证器和引理库最终是判断模型输出的语言模型。我们端到端地检测了这样一个系统,并分析了从被拒绝的参数中提取的 $51{,}754$ 跟踪观察结果($186$ 小时,\$$5{,}694$). We find three connected failure modes. First, the three-model verifier requires unanimity and treats parse or API failure as non-approval; in $10$ of $12$ verification events, one member returned no parseable output or an API error, making acceptance arithmetically impossible without surfacing an error. Second, when the ensemble did function, one verifier approved $3$ attempts that GPT rejected, each claiming to resolve the open problem; a single-verifier design would therefore have announced a solution three times. Third, because nothing could be approved, every review was a refutation, yet the lemma extractor mines reviews as well as proofs: $24$ of $93$ lemmas ($26\%$)。上下文已删除。总而言之,这些发现表明,如果没有外部验证,监督本身就是一个关键的信任边界:系统必须区分弃权和拒绝,保留有用的分歧,并在信息成为未来背景之前保留信息的来源和极性。

Viva La Vida:验证和多智能体证明搜索中的累积失败的原论文方法或结果图
图 3:Gemini 3.1 Pro 证明者声称 $\alpha^{\star}=3/5$ 的证明尝试。 (根跨度的显示总数(表 1 中报告的 $1,258.61) is Langfuse’s own aggregate; summing the per-generation costs over the same trace gives the $1,947。)