AuraTracer智迹闻
中文

EVENT DOSSIER

How many of your agent's calls actually need a frontier model?

NVIDIA released NeMo Switchyard and published test data on intelligent routing, showing that only 7% of tasks required advanced models, with costs reduced by 7…

2026-08-11 23:25 Models 🔥 31.9 heat score
2sources
1days unfolding
31.9heat score
2mentions
SummaryAI generated

NVIDIA has introduced NeMo Switchyard, designed to schedule AI Agent workloads across models through intelligent routing. In tests involving 145 agent tasks, it was found that only 7% of dialogue rounds required advanced models. Compared to sending all requests to the most powerful model, NeMo Switchyard improves accuracy by six points while reducing costs by 74%. This tool allows users to dynamically allocate different models based on task requirements, such as using classification models for specific steps, reasoning models for the next step, and small models for routine follow-up tasks, thereby optimizing model selection and resource allocation in the AI Agent construction process.

Related eventsRELATED EVENTS
Quick factsQUICK FACTS
145Number of tests
7%Proportion of advanced models used
Six pointsMultiple of accuracy improvement
74%Proportion of cost reduction
Key entitiesKEY ENTITIES
NVIDIANeMo Switchyard

Event frameEVENT FRAME

Launch

企业 · 官方发布across 1 days

Status

NVIDIA released NeMo Switchyard and published test data on intelligent routing, showing that only 7% of tasks required advanced models, with costs reduced by 7…

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
NVIDIA × NeMo Switchyard2

Integrated timelineUNIFIED TIMELINE

  1. 2026-08-11

    NVIDIA released NeMo Switchyard and test data

    NVIDIA introduced NeMo Switchyard for intelligent routing. In 145 task tests, only 7% required advanced models, with six points of accuracy improvement and 74% cost reduction.

    2 reports

SignalsSIGNALS

Keyword heat
  • NVIDIA2
  • NeMo Switchyard2

All reports (2)SOURCES

N NVIDIA Developer Blog en 2026-08-11 21:00

Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard

NVIDIA 推出 NeMo Switchyard,旨在跨模型调度 AI Agent 工作负载。该工具允许用户根据任务需求动态分配不同模型,以利用各模型的特定优势、弱点及成本特征。例如,单一代理任务可能需分类模型处理某一步骤、推理模型处理下一步,以及小模型处理常规跟进任务。若将所有请求发送至最大模型将导致效率低下或成本增加。NeMo Switchyard 通过智能路由解决此问题,优化 AI Agent 构建流程中的模型选择与资源分配。