IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
The Indian research team has developed a security assessment framework called IndicSafeEval, aimed at addressing the limitations of existing English-centered large language model security assessments. This framework was tested in four Indian languages: Hindi, Bengali, Marathi, and Punjabi, and generated 7,200 adversarial prompts covering ten categories of key content and security strategies. Using this framework, the researchers conducted black-box tests on open-source large language models and found significant differences in the security performance of models across languages and prompt styles, with varying levels of vulnerability for different risk categories. The results indicate that there is an urgent need to establish a multi-language and persuasion-perception benchmark framework to more accurately assess the security of large language models in real-world scenarios.