Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning
Under the zero-sample framework, this study utilized six years of outage records and high-resolution weather data to evaluate the risk prediction capability of large language models (LLMs) for weather-induced forced power outages in a utility service area in central Texas. The study defined the problem as a binary severity classification task across three forecast periods (3h, 6h, 12h), and compared the performance of four zero-sample LLMs with two supervised classifiers under two input configurations. The results showed that supervised models outperformed LLMs in terms of macro F1 score and accuracy rate, while the scores of the new generation LLMs were competitive. Additionally, LLMs provided complementary advantages in action-based reasoning and geographical scalability, indicating that combining LLMs with supervised models may be the best practice.