Corporate-Family Resolution Is Not a String-Matching Problem: A Public Benchmark Stratified by Name Visibility
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
The research team released the CorpFam benchmark dataset on September 7, 2026. This dataset incorporates 6,638,350 records from U.S. federal funding records, including 54,864 candidate pairs and 10,307 corporate families. The study evaluated the performance of the matcher by stratifying based on name visibility and found that the strongest matcher achieved a recall rate of 100.0% for identical pairs, while only 4.2% for invisible pairs. Since 93.2% of real links were never included in the candidate set due to filtering, the matching stage alone was insufficient to solve the problem of parsing corporate families. The study indicates that this task is essentially a retrieval problem rather than a simple string matching issue, and the key intervention point lies in the candidate generation stage rather than the ranking stage. The team has made the benchmark dataset, judgment logs, and source code available for subsequent research.