AuraTracer智迹闻
中文

EVENT DOSSIER

Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs

2026-09-07 12:00 Science 🔥 40.2 heat score
1sources
1days unfolding
40.2heat score
0mentions
SummaryAI generated

On September 7, 2026, arXiv cs.CL published the paper “Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs”. This study aims to process data of poetry and commentary pairs in classical Tamil using representation learning techniques.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.CL en 2026-09-07 12:00

Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs

研究人员构建了包含 1,262 对古典泰米尔语诗歌 - 注释(urai)语料库,涵盖从技术语法到现代释义的五个来源部分。研究训练了循环神经网络、Transformer 编码器、孪生风格配对匹配网络、mBART 风格编解码器及仅解码器语言模型,并针对同一数据设置了相应控制组进行对比分析。结果显示,包含 25 个最频繁注释词的固定字符串在生成重叠度上优于仅解码器模型;在此样本规模下,正态噪声上的规范相关系数达 1.000,而 token-F1 分数仅为 0.02-0.20,且编解码器在验证损失上升后仍持续降低训练损失十六个周期。唯一显著结果为:仅解码器模型在 112 对最小配对比较中偏好真实词序(107 例,95.5%),但无法复现未参与训练的注释内容。研究团队已发布提取与评估协议,原始注释的再分发需经许可。