On August 24, 2026, MIT Tech Review AI reported that human children still outperform the most complex language models significantly in mastering their native languages. Although LLMs like ChatGPT, Claude, and DeepSeek can engage in fluent conversations, their training requires vast amounts of data, while children only need about 100 million to 300 million words to become proficient in language. Michael C. Frank, a cognitive scientist at Stanford University, pointed out that LLMs need to process corpus containing the experience of an entire generation in a city to achieve the milestones that children can accomplish within a year. As the availability of data on the Internet may become scarce, creating more efficient AI models by reverse-engineering human learning mechanisms becomes a key challenge, which helps to address scientific questions such as the nature of language acquisition and child mental development.
“Human children still outperform the most complex language models in mastering their native language, a phenomenon known as the ‘data efficiency gap’. Although LLMs such as ChatGPT, Claude, DeepSeek, and GPT can engage in fluent conversations, their training requires vast amounts of data, while children only need about 100 to 300 million words to become proficient in language. Michael C. Frank, a cognitive scientist at Stanford University, points out that LLMs need to process corpus containing the experience of an entire generation in a city to achieve the milestones that children accomplish within a year. As access to data through the Internet may become scarce, creating more efficient AI models by reverse-engineering human learning mechanisms becomes a key challenge, which can help address scientific questions such as the nature of language acquisition and child mental development.”