Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
The KoNA benchmark was released, revealing that visual language models have selective non-compliance flaws in specific scenarios.
2026-09-07 12:00Models🔥 47.2 heat score
2sources
1days unfolding
47.2heat score
5mentions
SummaryAI generated
On September 7, 2026, researchers released the KoNA benchmark, aimed at evaluating the selective non-compliance capabilities of visual language models in five scenarios: false premises, visual accessibility issues, general unknowns, task feasibility, and security. The benchmark included 3,100 instances, and comparative tests using single sentences and complex queries revealed that existing models often failed to properly reject or remain silent, especially in complex queries where selective non-compliance was required. To address this, researchers used KoNA data containing examples of selective non-compliance and fully answerable tasks for fine-tuning. The results showed that this approach significantly improved the accuracy of non-compliance responses while maintaining the model’s performance on fully answerable tasks, enabling it to distinguish between answerable components and those requiring non-compliance and respond appropriately.
Hugging FaceHyounghun KimJihyoung JangKoNAMinji Kim
Event frameEVENT FRAME
Research
政府 · 科研机构across 1 days
Status
The KoNA benchmark was released, revealing that visual language models have selective non-compliance flaws in specific scenarios.
Coverage · reports per dayLANGUAGE SPLIT
Entity relations
Integrated timelineUNIFIED TIMELINE
2026-09-07
KoNA benchmark released
Researchers released the KoNA benchmark to evaluate the selective non-compliance capabilities of visual language models in five scenarios: false premises, visual accessibility issues, general unknowns, task feasibility, and security. This benchmark includes 3,100 instances, and comparative tests using single sentences and compound queries reveal that existing models often fail to properly reject or remain silent.
Minji Kim 等人在论文《Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models》中提出 KoNA 基准,用于评估视觉语言模型在假前提、视觉不可访问性、普遍未知、任务不可行及不安全等五类场景下的选择性不合规能力。该基准包含 3,100 个实例,通过单问与复合问对进行对比测试。现有研究表明,模型往往无法正确拒绝或保持沉默,且在需要选择性不合规的复合问中表现更差。作者提出使用 KoNA 示例配合完全可回答数据集对视觉语言模型进行微调,该方法在显著提升不合规准确性的同时,基本保持了完全可回答查询及通用基准的性能。
研究人员发布 KoNA 基准,用于评估视觉语言模型在五个类别(错误前提、视觉不可访问性、通用未知、任务可行性、安全)中的选择性不合规能力。该基准通过配对单句和复合查询测试查询级与组件级不合规能力,发现现有模型常无法恰当拒绝或保持沉默,且此类失败在选择性不合规场景下更为显著。为此,研究者利用包含选择性不合规示例及完全可回答任务的 KoNA 数据进行微调,结果显示微调后的模型在不合规准确性上取得显著提升,同时基本保持了在完全可回答任务上的表现,能够区分可回答组件与需不合规的组件并以恰当方式响应。