AuraTracer智迹闻
中文

EVENT DOSSIER

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated

On September 7, 2026, the Harbor team released Harbor Adapters and Harbor-Index, aimed at building a large-scale Agent evaluation system. Harbor Adapters provide a unified infrastructure that supports adapting over 80 benchmarks to any Agent evaluation, and the effectiveness of Benchmark Adapters was verified through code review and peer experiments. The research team conducted large-scale evaluations on 8 models covering 54 benchmarks, all of which were run using Terminus-2 and three native Harness versions. Additionally, the team introduced Harbor-Index, a curated meta-data set containing 82 challenging and diverse tasks. This dataset was derived from suitable templates, filtered by difficulty, refined through AI and manual auditing, and subjected to a “audit-repair” cycle, maintaining the challenge and breadth of large-scale evaluations while keeping operating costs manageable. In all models…

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
CodexGPT-5.5Harbor AdaptersHarbor-IndexTerminus-2

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Codex × GPT-5.51Codex × Harbor Adapters1Codex × Harbor-Index1Codex × Terminus-21GPT-5.5 × Harbor Adapte…1GPT-5.5 × Harbor-Index1

SignalsSIGNALS

Keyword heat
  • Harbor Adapters1
  • Harbor-Index1
  • Terminus-21
  • GPT-5.51
  • Codex1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Harbor Adapters has released a unified evaluation infrastructure, supporting the adaptation of over 80 benchmarks for any Agent evaluation. The research team developed Benchmark Adapters and verified them through code review and peer experiments; large-scale evaluations were conducted on 8 models covering 54 benchmarks, each run using Terminus-2 and 3 native Harness configurations. Additionally, the team introduced Harbor-Index, a carefully selected meta-data set containing 82 challenging and diverse tasks, derived from a suitable set of datasets filtered by difficulty, audited by AI and manually, and refined through a “audit-repair” cycle. This dataset maintains the challenge and breadth of large-scale evaluations while keeping operating costs manageable; the pass rate for all model-Harness configurations does not exceed 30%, with the strongest combination (GPT-5.5 with Codex) reaching 28.0%. Related adapters, evaluation results, in-depth analysis, and Harbor-…