AuraTracer智迹闻
中文

EVENT DOSSIER

RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
4mentions
SummaryAI generated

The `RealCADBench benchmark set has been officially released, aiming to comprehensively evaluate the parametric computer-aided design (CAD) modeling capabilities based on real industrial design intentions. This dataset includes 12,632 tasks from 19 categories of factory automation, covering various input modalities such as text descriptions, 2D engineering drawings, real product images, and rendered images, supporting modeling of parts and assemblies. In the test consisting of 1,770 evaluation slices, models were verified by generating FreeCAD API Python code and exporting 3D models during shared runtime. The evaluation metrics include executability, solid IoU, surface IoU, and rule-based visual semantic identity detection. The results show that none of the nine independent advanced models currently available leads comprehensively in any of the four core metrics; the executability scores of six advanced-scale models range from 0.565 to 0.812, the solid IoU range from 0.2841 to 0.5379, and the surface IoU ranges from 0.11…

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
CodexFreeCADGPT-5.5RealCADBench

Event frameEVENT FRAME

Launch

arXiv:2609.03773v2 RealCADBench 发布工业参数化 CAD 建模基准

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Codex × FreeCAD1Codex × GPT-5.51Codex × RealCADBench1FreeCAD × GPT-5.51FreeCAD × RealCADBench1GPT-5.5 × RealCADBench1

SignalsSIGNALS

Keyword heat
  • RealCADBench1
  • FreeCAD1
  • GPT-5.51
  • Codex1

All reports (1)SOURCES

A arXiv cs.CL en 2026-09-07 12:00

RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

RealCADBench 发布,旨在评估基于真实工业设计意图的参数化 CAD 建模。该基准包含来自 19 个工厂自动化类别的 12,632 项任务,涵盖文本描述、2D 工程图纸、真实产品图片及渲染图像,支持 Part 和 Assembly 建模。在 1,770 项评估切片中,各方法生成 FreeCAD API Python 代码并通过共享运行时导出 3D 模型,使用可执行性、Solid IoU、Surface IoU 及基于规则的视觉语义身份 Judge 进行评估。结果显示,九种独立前沿大模型中无单一模型在所有四项指标上领先;六种前沿规模大模型的可执行性范围在 0.565 至 0.812 之间,Solid IoU 在 0.2841 至 0.5379 之间,Surface IoU 在 0.112 至 0.217…