VERGE: Verification-Enhanced Refinement for Grounded Extraction of Early-Onset Colorectal Cancer Symptoms in Clinical Notes
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
The researchers developed a proxy workflow called VERGE, designed to automatically extract six red-flag symptoms of early colorectal cancer and family history risk status from clinical notes. The model uses retrieval-enhanced generation techniques to generate initial labels and evidence, followed by a controlled verification and refinement cycle to examine textual basis and clinical effectiveness; when limitations or unsolvable problems are encountered, they are handed over for manual review. On a test set containing 4,033 pairs of annotated notes from doctors, VERGE performed better than single-agent baselines, rule-based clinical language processing baselines, and other alternative language models. Specifically, compared with the single-agent baseline, VERGE effectively reduced the false positive rate, improving accuracy from 0.764 to 0.849, and Matthews correlation coefficient (MCC) from 0.681 to 0.730. Additionally, the model was able to independently resolve most labeling errors, with only 1.5% of statements requiring manual review.
The researchers developed a proxy workflow called VERGE, designed to automatically extract six red flag symptoms of early colorectal cancer and family history risk status from clinical notes. VERGE proposes initial labels and evidence through enhanced retrieval, then enters a controlled verification-refinement cycle to examine textual basis and clinical validity. When limitations are reached or issues cannot be resolved, it transfers cases to manual review. The study evaluated the system on a test set containing 4,033 pairs of annotated doctor notes, comparing it with a single-proxy baseline, a rule-based clinical language processing baseline, and alternative language models. Compared to the single-proxy baseline, VERGE reduced the false positives rate, improved accuracy from 0.764 to 0.849, and increased the Matthews correlation coefficient (MCC) from 0.681 to 0.730. It also independently resolved most labeling errors, with only 1.5% of statements requiring manual review.