HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
The HarvestBench research paper has been published, showing significant differences among nine models in animal avoidance tasks.
2026-09-07 12:00Models🔥 47.2 heat score
2sources
1days unfolding
47.2heat score
4mentions
SummaryAI generated
HarvestBench is the first benchmark that aims to avoid harming living animals in pricing decisions, proposed by Jasmine Brazilek and others. The study simulates the decision-making behavior of LLM agents driving tractors during corn harvesting in farmland, covering 7,201 pricing decisions across nine models. Experiments revealed significant differences in the killing rates of various models, ranging from 0.4% to 98.8%. Terra and Sol were the most humane, while GPT-4o-mini was the most cruel, and its performance had nothing to do with the model’s capabilities. Under default settings, all models more frequently crushed wild animals than livestock. After introducing a moral briefing, the killing rate of five-sixths of the reasoning models dropped below 6%; removing this briefing increased the killing rate to over 84%. Additionally, three-quarters of the models were sensitive to prices at 5%. The benchmark uses game log counts for scoring, without the need for an LLM evaluator to ensure reproducibility.
The HarvestBench research paper has been published, showing significant differences among nine models in animal avoidance tasks.
Coverage · reports per dayLANGUAGE SPLIT
Entity relations
Integrated timelineUNIFIED TIMELINE
2026-09-07
HarvestBench research paper published
arXiv cs.AI and Hugging Face papers published HarvestBench benchmark results, simulating LLM agents harvesting corn while facing animal decisions. It covers 7,201 pricing decisions across nine models, with bacterial killing rates ranging from 0.4% to 98.8%.
HarvestBench is the first benchmark that aims to avoid harming living animals in its pricing criteria. It measures the moral performance of models by simulating their decision-making behavior when LLM agents drive tractors to harvest corn. The study ran nine models through 7,201 pricing decisions in a reinforcement learning environment with animals, and found that the rate of killing animals ranged from 0.4% to 98.8%. Terra and Sol were the most benevolent, while GPT-4o-mini was the most cruel. The results showed that all models more frequently crushed wild animals rather than domestic livestock under the default map; after introducing a moral briefing, the killing rate of five out of six inference models dropped below 6%, but removing the briefing caused the killing rate to soar above 84%. Additionally, three-quarters of the models were sensitive to a price of 5%, and the benchmark uses game log counts for scoring, without the need for an LLM evaluator to ensure reproducibility.