A Removal Based Approach to Improve LLM Faithfulness at Test-Time
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
To address the issue of incomplete explanations for Large Language Models (LLMs), researchers proposed a new method based on removal. This method is implemented during the testing phase by removing concepts that are not referenced by the model and re-questioning the model, aiming to eliminate the influence of unmentioned factors while retaining the impact of mentioned concepts. Experiments show that this method outperforms standard prompts and existing methods for encouraging fidelity on two datasets, various model families, and two independent fidelity metrics. This method is model-independent and can be applied during reasoning without modifying model parameters, providing a flexible mechanism for improving the reliability and safety of decision-making assisted by LLMs.
This paper proposes a new method to directly address the issue of incomplete explanations in large language models (LLMs) during the testing phase by removing concepts not mentioned in the model explanations. This method eliminates unmentioned factors while retaining the influence of mentioned concepts by removing unexplainably referenced concepts from the input and re-questioning the model. Experiments show that this method outperforms standard prompts and prompts designed to encourage fidelity on two datasets, various model families, and two independent fidelity metrics. This method is model-independent and can be applied during reasoning without modifying model parameters, providing a flexible mechanism for reducing hidden influences and enhancing the reliability and security of LLM-assisted decision-making.