AuraTracer智迹闻
中文

EVENT DOSSIER

KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
8mentions
SummaryAI generated

On September 7, 2026, arXiv released KernelGenBench, the first unified multi-source, multi-chip infrastructure for evaluating large language models and agent-generated Triton kernels. This benchmark covers six hardware platforms, including 210 operators from PyTorch ATen, production vLLM, and proprietary cuBLAS, with 110 operators tested to be semantically stable on all six platforms. The evaluation process consumed over 15 billion tokens. The results showed that while agent execution improved accuracy, none of the methods had a clear advantage across all sources and platforms: vLLM presented the strongest challenge in accuracy, cuBLAS set the upper limit for performance, and AutoKernel’s accuracy dropped from 87% on the NVIDIA platform to 25% on the Iluvatar CoreX platform. The dedicated agents averaged approximately 4.99 million tokens per successful operator.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
AutoKernelIluvatar CoreXKernelGenBenchNVIDIAPyTorch ATenTritoncuBLASvLLM

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
AutoKernel × Iluvatar C…1AutoKernel × KernelGenB…1AutoKernel × NVIDIA1AutoKernel × PyTorch AT…1AutoKernel × Triton1Iluvatar CoreX × Kernel…1

SignalsSIGNALS

Keyword heat
  • KernelGenBench1
  • Triton1
  • PyTorch ATen1
  • vLLM1
  • cuBLAS1
  • AutoKernel1
  • NVIDIA1
  • Iluvatar CoreX1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation

arXiv:2607.27231v2 发布 KernelGenBench,首个统一多源多芯片基础设施用于评估 LLM 及智能体生成的 Triton 内核。该基准覆盖六款硬件平台,包含来自 PyTorch ATen、生产 vLLM 及专有 cuBLAS 的 210 个算子(KernelGenBench-MS),并在六款平台上测试语义稳定的 110 个算子(KernelGenBench-MC)。评估消耗超过 150 亿 token,结果显示智能体执行提升了正确性,但无方法在所有源和平台上占优:vLLM 带来最强正确性挑战,cuBLAS 设定最高性能上限,AutoKernel 准确率从 NVIDIA 的 87% 降至 Iluvatar CoreX 的 25%。专用智能体平均每个成功算子消耗 499 万 token,…