AuraTracer智迹闻
中文

EVENT DOSSIER

An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

The researchers proposed an integrated model called ChatNDE Figure to Caption, aimed at automating image analysis for non-destructive testing (NDE) images using deep learning and natural language processing techniques. This system combines ResNet50 for extracting image features with GPT2 for generating text, and integrates a visual question answering (VQA) module to answer questions such as “Is there a crack?” The research team constructed a large NDE dataset and evaluated the model’s output against expert descriptions using BLEU scores. Currently, the model’s overall accuracy is low, but the generated brief descriptions can cover the key features of images that human experts focus on. This approach helps to accelerate the NDE testing process, reduce human errors, and improve technical accessibility.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
ChatNDE Figure to CaptionGPT2ResNet50

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
ChatNDE Figure to Capti…1ChatNDE Figure to Capti…1GPT2 × ResNet501

SignalsSIGNALS

Keyword heat
  • ChatNDE Figure to Caption1
  • ResNet501
  • GPT21

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis

引入名为 ChatNDE Figure to Caption 的 AI 方法,旨在利用深度学习和自然语言处理自动化无损检测(NDE)图像分析。该系统结合 ResNet50 提取图像特征与 GPT2 生成自然文本,并集成视觉问答(VQA)模型以回答如“是否存在裂纹”等具体问题。研究构建了大型 NDE 数据集,通过 BLEU 分数评估模型输出与专家描述的匹配度。尽管当前模型准确率较低,但生成的简短 caption 已能涵盖人类专家关注的图像关键特征。该方案有助于加速 NDE 工作流程、减少人为错误并提升技术可及性。