论文

READ-Bench:时间序列诊断的历史实例检索基准测试

READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis

模型评测模型能力评测

摘要

Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so visually different traces of the same fault count as relevant while similar-looking traces of different faults do not.将检索视为基础检索器,然后是重排序器,我们在一种通过显着性测试改变监督、污染和语料库规模的协议下评估经典距离、符号检索器、自监督和基础模型嵌入器及其融合,加上标签感知和语言模型重排序器。在通用的独立于通道的接口下,与单独搜索的强大的经典和符号基线相比,预训练的表示没有提供统计上可检测的优势。决定性因素是重排序时的少量已解决案例监督,即高斯过程重排序器,它在嵌入空间中传播一些邻居标签,这比更复杂的表示或语言模型推理更有帮助,并且在污染和全语料库规模下保持不变。 Guided by these findings, we fuse a normal-residual-scored embedder with a dynamic time warping leg via reciprocal-rank fusion, then rerank with the Gaussian-process reranker, improving NDCG@10 over its own search stage on all 12 datasets, by +0.11 from reranking and +0.16 over the strongest single base retriever.

READ-Bench:时间序列诊断的历史实例检索基准测试:图 6:READ-Bench 数据集中的代表性异常类。每一行都是一个数据集,每个单元格显示一种故障类型(红色)的代表性窗口与正常窗口(灰色)的对比。
图 6:READ-Bench 数据集中的代表性异常类。每一行都是一个数据集,每个单元格显示一种故障类型(红色)的代表性窗口与正常窗口(灰色)的对比。