AuraTracer智迹闻
中文

EVENT DOSSIER

Orchard: An open framework for scalable agentic AI

2026-08-04 00:00 Models 🔥 28.9 heat score
1sources
1days unfolding
28.9heat score
5mentions
SummaryAI generated

On August 3, 2026, Microsoft Research released an open-source framework named Orchard, aimed at creating a scalable and cost-effective research environment for agent-based AI. The core of this framework is Orchard Env, which supports the training and evaluation of various agents such as software, web navigation, and personal assistants. It can be directly run in real-deployment environments like Codex and OpenClaw. Experiments show that models with only about 3 billion parameters can achieve an accuracy of 69.7% on the SWE-bench Verified benchmark (73.0% after reordering), which is comparable to systems using over ten times more parameters. The project also makes training data, evaluation methods, and domain-specific training recipes available, aiming to address the bottleneck in agent-based research caused by reliance on private infrastructure.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
CodexOpenClawOrchardOrchard EnvZeroClaw

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Codex × OpenClaw1Codex × Orchard1Codex × Orchard Env1Codex × ZeroClaw1OpenClaw × Orchard1OpenClaw × Orchard Env1

SignalsSIGNALS

Keyword heat
  • Orchard1
  • Orchard Env1
  • Codex1
  • OpenClaw1
  • ZeroClaw1

All reports (1)SOURCES

M Microsoft Research en 2026-08-04 00:00

Orchard: An open framework for scalable agentic AI

Orchard has released an open-source framework aimed at creating a scalable and cost-effective environment for agent-based AI research. The core of this framework is Orchard Env, which supports training and evaluation of various agents, including software, web navigation, and personal assistants. It can be directly run in real-deployment environments such as Codex and OpenClaw. Experiments show that models with only approximately 3 billion parameters can achieve an accuracy of 69.7% on SWE-bench Verified (73.0% after reordering), which is close to the performance of systems using ten times more parameters. The project also makes training data, evaluation methods, and domain-specific training recipes available, aiming to address the bottleneck in current agent research caused by reliance on private infrastructure.