AuraTracer智迹闻
中文

EVENT DOSSIER

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

The Indian research team has developed a security assessment framework called IndicSafeEval, aimed at addressing the limitations of existing English-centered large language model security assessments. This framework was tested in four Indian languages: Hindi, Bengali, Marathi, and Punjabi, and generated 7,200 adversarial prompts covering ten categories of key content and security strategies. Using this framework, the researchers conducted black-box tests on open-source large language models and found significant differences in the security performance of models across languages and prompt styles, with varying levels of vulnerability for different risk categories. The results indicate that there is an urgent need to establish a multi-language and persuasion-perception benchmark framework to more accurately assess the security of large language models in real-world scenarios.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
IndicSafeEval

SignalsSIGNALS

Keyword heat
  • IndicSafeEval1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

印度研究团队推出 IndicSafeEval 框架,针对印地语、孟加拉语、马拉地语和旁遮普语四种印度语言开展安全鲁棒性评估,生成涵盖十类关键内容与安全策略的 7,200 个对抗提示。该框架对开源大语言模型进行黑盒测试,发现模型在不同语言和提示风格下的安全表现存在显著差异,且不同风险类别的脆弱程度不一。现有以英语为中心的安全评估方法存在局限,亟需建立多语言及说服感知的基准框架以更准确评估真实世界中的大语言模型安全性。