FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders
This paper proposes the FailSAE framework, which uses sparse autoencoders to achieve interpretable fault prediction for visual language models. This method defines fault prediction as a classification task for the latent activations of sparse SAEs, and introduces a three-stage fault-aware training process aimed at enhancing the amount of fault prediction information while maintaining interpretability. Experiments show that this framework outperforms existing baseline methods in fault prediction performance. Further analysis reveals that fault-aware training enables SAE latent directions to capture more category-specific concepts; it also shows how the model representation shifts from category-specific concepts to vague or style-related concepts during faults. Additionally, the study explores the supporting role of the learned SAE latent directions in fault recovery during operation.