Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
The researchers replicated and expanded upon Wu et al. (2026)’s sparse autoencoder (SAE) feature discovery method in the Gemma 2 and Gemma 3 models, aiming to verify the causal effect of translation initiation features in multilingual environments. The test results showed that although both models identified more than 20 features frequently activated across settings, causal validation indicated that these features had little or no consistent impact on translation behavior. Only one feature, (L10, 5717) in Gemma 2 and (L20, 2456) in Gemma 3, exhibited consistent performance across 23 language settings: amplifying its activation improved the COMET score, while suppressing it decreased the score. The results suggest that feature repetition overestimates cross-language transfer capabilities and identified a common translation initiation direction in both Gemma 2 and Gemma 3.