From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
To address the issue of unstable capability estimation in existing large language model routing methods due to the randomness of single-sample responses, researchers proposed a new method called Distribution-Aware Routing Supervision (DARS). This method estimates the distribution of model capabilities at the query level through semantic-preserving query rewriting and multiple random decoding observations. It combines expected quality, cost, and performance fluctuations to construct risk-aware supervision without changing the downstream routing architecture. Experiments across various tasks and routing methods show that DARS improves routing efficiency and optimizes the cost-quality balance compared to single-sample supervision. Further analysis indicates that this advantage remains effective with medium sampling budgets and different decoding temperatures.