On September 7, 2026, arXiv cs.AI published the paper “Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods”. The study proposes that in reference-based automatic evaluation methods, “behavioral correctness assumptions” should be introduced to go beyond traditional aggregate score evaluations.