AuraTracer智迹闻
中文

EVENT DOSSIER

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, arXiv released PRISM-Bench, the first audio-centric diagnostic benchmark for text-generation audio and video (T2AV). This benchmark is built based on 900 manually verified samples and divides audio evaluation into two orthogonal dimensions: audio type (voice, music, sound) and sound source visibility (inside the screen, outside the screen). It evaluates from four perceptual dimensions: audiovisual consistency, audio quality, audio expressiveness, and prompt adherence, comprising a total of 35 detailed criteria. To ensure the reliability of the evaluation, the study used an enhanced multimodal large model based on blind side-by-side comparison as the evaluation protocol, with an average一致性 of over 70% with human scorers. Recent evaluations of T2AV systems show that there is a significant performance gap between cutting-edge models and open-source models; it was also found that the current generation paradigm overfits perceptual fidelity and performs poorly in complex localization and control tasks, especially in music generation and screen-in audio synchronization.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
PRISM-Bench

SignalsSIGNALS

Keyword heat
  • PRISM-Bench1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

PRISM-Bench 发布了首个面向文本生成音视频(T2AV)的音频中心诊断基准。该基准基于 900 个经人工验证的样本构建,将音频评估分解为音频类型(语音、音乐、声音)和声源可见性(屏幕内、屏幕外)两个正交维度,并从视听一致性、音频质量、音频表现力和提示遵循四个感知维度进行评价,共包含 35 项细粒度标准。为确保评估可靠性,研究采用了基于盲测并排对比的增强型多模态大模型作为裁判协议,其与人类评分者的平均一致性超过 70%。对近期 T2AV 系统的评估显示,前沿模型与开源模型之间存在显著性能差距;同时发现当前生成范式过度拟合感知保真度,而在复杂定位与控制任务上表现不佳,特别是在音乐生成和屏幕内音频同步方面。