SciTech Pulse
AI

Why Children Still Out-Learn AI at Language, and What It Could Teach Machines

Human children learn to speak fluently on a tiny fraction of the words that large language models need, a gap researchers call the data efficiency gap -- and scientists hope reverse-engineering it could lead to more…

Large language models (LLMs), the AI systems behind chatbots such as Claude, DeepSeek and OpenAI's GPT models, can converse fluently -- but only after training on far more words than any human ever hears. Meta's open-weight LLM Llama 3.1, released two years ago, was trained on 15 trillion tokens, or word-like chunks of language, and Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University, says frontier models could be pretrained on ten times more data than that.

A human child, by contrast, typically starts producing grammatically correct sentences after hearing something like 10 million to 30 million words. A preteen raised in a language-rich home may have heard around 100 million words, or up to 300 million once reading is added, by age 20. Researchers call this gap between children and machines the data efficiency gap. "Claude has seen the amount of language that an entire city will experience in one generation," Wilcox said.

Michael C. Frank, a cognitive scientist at Stanford University, said that despite recent progress in AI, "we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year." He added that training a model like GPT-2 on the roughly 30 million words a child hears produces "a nonsense generator," not a child.

Scientists including Frank and Wilcox are trying to reverse-engineer how children learn language with so little data, hoping the insights could lead to more data-efficient AI models -- useful for training AI on video or building chatbots for minority languages -- and could also help settle long-running debates about whether humans are born with innate knowledge of grammar, an idea proposed by linguist Noam Chomsky, or learn language purely through experience, as psychologist B.F. Skinner argued.

Terms explained

The story so far

  1. AI-Designed 'Intrabodies' Could Open New Paths to Treating Alzheimer's, Parkinson's and MND
  2. AI May Know How You'll Respond to a Vaccine Before You Get It
  3. When AI Designs a Drug, the Patent Still Names a Human
  4. New Mathematical Framework Could Cut AI Memory's Energy Use by Thousands of Times
  5. AI Mapping Reveals a Hidden Stage of Arctic Freeze, With Climate Implications
  6. Bangladesh's Team Atlas Wins Global Robotics Title in Japan
  7. Why Children Still Out-Learn AI at Language, and What It Could Teach Machines
#AI#language models#cognitive science#language acquisition
Rate this story

Related stories