AuraTracer智迹闻
中文

EVENT DOSSIER

NxN E-valuation: Hypothesis Certification via a Conformal CRT Null

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, arXiv published the paper “NxN E-valuation: Hypothesis Certification via a Conformal CRT Null” (arXiv:2608.06621v3), introducing a new algorithm called NxN E-valuation. This method relies on e-value for hypothesis verification, aiming to address the issue of幻觉 produced by large language models (LLMs). Unlike existing methods, it does not require constructing certification procedures or dedicated null hypotheses for specific cases; instead, it utilizes a sufficiently large natural training dataset to use different samples as null hypotheses in conditional randomization tests (CRT), thereby directly verifying each hypothesis. This algorithm is designed specifically for LLM exploration and can serve as a general alternative to cyclic verification and holdout testing, provided that the hypotheses generated by the LLM are applicable to each independent sample.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
NxN E-valuation

SignalsSIGNALS

Keyword heat
  • NxN E-valuation1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

NxN E-valuation: Hypothesis Certification via a Conformal CRT Null

arXiv:2608.06621v3 提出 NxN E-valuation,这是一种基于 e-value 的假设验证算法,无需构建特定案例的认证程序(如专用零假设),仅需足够大的数据集即可验证假设。该方法专为大语言模型(LLM)探索系统设计,旨在解决 LLM 虽擅长提出假设但易产生幻觉的问题,以替代现有的循环验证和留样测试等不足方案。NxN E-valuation 利用自然存在的大训练集,让不同样本互为零假设来执行条件随机化检验(CRT),从而直接认证每个假设。该方法可作为循环验证和留样数据的通用更好替代品,前提是 LLM 生成的假设适用于每个独立样本。