AuraTracer智迹闻
中文

EVENT DOSSIER

FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

The researchers proposed the FailSAE framework, using sparse autoencoders to solve the fault prediction problem in visual language models. This method defines fault prediction as a classification task for the potential activations of SAEs, and introduces a three-stage fault-aware training process aimed at enhancing the interpretability and information content of the potential directions. Experiments show that this framework performs better than existing baseline methods. Analysis indicates that fault-aware training enables SAEs to capture more category-specific concepts, while revealing the representational changes in the model during faults, shifting from category-specific concepts to vague or style-related concepts. Additionally, the study explored the support of the learned SAE potential directions for fault recovery during operation.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
CLIP

SignalsSIGNALS

Keyword heat
  • CLIP1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders

This paper proposes the FailSAE framework, which uses sparse autoencoders to achieve interpretable fault prediction for visual language models. This method defines fault prediction as a classification task for the latent activations of sparse SAEs, and introduces a three-stage fault-aware training process aimed at enhancing the amount of fault prediction information while maintaining interpretability. Experiments show that this framework outperforms existing baseline methods in fault prediction performance. Further analysis reveals that fault-aware training enables SAE latent directions to capture more category-specific concepts; it also shows how the model representation shifts from category-specific concepts to vague or style-related concepts during faults. Additionally, the study explores the supporting role of the learned SAE latent directions in fault recovery during operation.