Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning
A new study suggests that shared circuits can predict the ability of large language models (LLMs) to generalize across formats in arithmetic reasoning. The researchers used attribution patch technology to independently identify the different circuits used by the model to solve numerical arithmetic and text-based problems in English, Spanish, and Italian. They tested whether this overlap could predict the model’s generalization performance to text formats. The study found that circuit overlap can explain the differences in relative difficulty among the three text formats and accurately predict the model’s generalization performance and correct solution rates, with results comparable to those achieved by supervised probing but without the need for labeled data.