AuraTracer智迹闻
中文

EVENT DOSSIER

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

The KoNA benchmark was released, revealing that visual language models have selective non-compliance flaws in specific scenarios.

2026-09-07 12:00 Models 🔥 47.2 heat score
2sources
1days unfolding
47.2heat score
5mentions
SummaryAI generated

On September 7, 2026, researchers released the KoNA benchmark, aimed at evaluating the selective non-compliance capabilities of visual language models in five scenarios: false premises, visual accessibility issues, general unknowns, task feasibility, and security. The benchmark included 3,100 instances, and comparative tests using single sentences and complex queries revealed that existing models often failed to properly reject or remain silent, especially in complex queries where selective non-compliance was required. To address this, researchers used KoNA data containing examples of selective non-compliance and fully answerable tasks for fine-tuning. The results showed that this approach significantly improved the accuracy of non-compliance responses while maintaining the model’s performance on fully answerable tasks, enabling it to distinguish between answerable components and those requiring non-compliance and respond appropriately.

Related eventsRELATED EVENTS
Quick factsQUICK FACTS
3,100Number of instances
5Number of evaluation categories
Key entitiesKEY ENTITIES
Hugging FaceHyounghun KimJihyoung JangKoNAMinji Kim

Event frameEVENT FRAME

Research

政府 · 科研机构across 1 days

Status

The KoNA benchmark was released, revealing that visual language models have selective non-compliance flaws in specific scenarios.

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Hugging Face × Hyounghu…1Hugging Face × Jihyoung…1Hugging Face × Minji Kim1Hyounghun Kim × Jihyoun…1Hyounghun Kim × Minji K…1Jihyoung Jang × Minji K…1

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-07

    KoNA benchmark released

    Researchers released the KoNA benchmark to evaluate the selective non-compliance capabilities of visual language models in five scenarios: false premises, visual accessibility issues, general unknowns, task feasibility, and security. This benchmark includes 3,100 instances, and comparative tests using single sentences and compound queries reveal that existing models often fail to properly reject or remain silent.

    2 reports

SignalsSIGNALS

Keyword heat
  • KoNA1
  • Hugging Face1
  • Minji Kim1
  • Jihyoung Jang1
  • Hyounghun Kim1

All reports (2)SOURCES

H Hugging Face Papers en 2026-09-07 08:00

Paper page - Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

Minji Kim 等人在论文《Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models》中提出 KoNA 基准,用于评估视觉语言模型在假前提、视觉不可访问性、普遍未知、任务不可行及不安全等五类场景下的选择性不合规能力。该基准包含 3,100 个实例,通过单问与复合问对进行对比测试。现有研究表明,模型往往无法正确拒绝或保持沉默,且在需要选择性不合规的复合问中表现更差。作者提出使用 KoNA 示例配合完全可回答数据集对视觉语言模型进行微调,该方法在显著提升不合规准确性的同时,基本保持了完全可回答查询及通用基准的性能。

A arXiv cs.AI en 2026-09-07 12:00

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

研究人员发布 KoNA 基准,用于评估视觉语言模型在五个类别(错误前提、视觉不可访问性、通用未知、任务可行性、安全)中的选择性不合规能力。该基准通过配对单句和复合查询测试查询级与组件级不合规能力,发现现有模型常无法恰当拒绝或保持沉默,且此类失败在选择性不合规场景下更为显著。为此,研究者利用包含选择性不合规示例及完全可回答任务的 KoNA 数据进行微调,结果显示微调后的模型在不合规准确性上取得显著提升,同时基本保持了在完全可回答任务上的表现,能够区分可回答组件与需不合规的组件并以恰当方式响应。