SciTech Pulse
AI

Why Children Still Out-Learn AI at Language, and What It Could Teach Machines

Human children learn to speak fluently on a tiny fraction of the words that large language models need, a gap researchers call the data efficiency gap -- and scientists hope reverse-engineering it could lead to more…

Large language models (LLMs), the AI systems behind chatbots such as Claude, DeepSeek and OpenAI's GPT models, can converse fluently -- but only after training on far more words than any human ever hears. Meta's open-weight LLM Llama 3.1, released two years ago, was trained on 15 trillion tokens, or word-like chunks of language, and Ethan Gotlieb Wilcox, a cognitive scientist and linguist at Georgetown University, says frontier models could be pretrained on ten times more data than that.

A human child, by contrast, typically starts producing grammatically correct sentences after hearing something like 10 million to 30 million words. A preteen raised in a language-rich home may have heard around 100 million words, or up to 300 million once reading is added, by age 20. Researchers call this gap between children and machines the . "Claude has seen the amount of language that an entire city will experience in one generation," Wilcox said.

Michael C. Frank, a cognitive scientist at Stanford University, said that despite recent progress in AI, "we still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year." He added that training a model like GPT-2 on the roughly 30 million words a child hears produces "a nonsense generator," not a child.

Scientists including Frank and Wilcox are trying to reverse-engineer how children learn language with so little data, hoping the insights could lead to more data-efficient AI models -- useful for training AI on video or building chatbots for minority languages -- and could also help settle long-running debates about whether humans are born with innate knowledge of grammar, an idea proposed by linguist Noam Chomsky, or learn language purely through experience, as psychologist B.F. Skinner argued.

Terms explained

The story so far

  1. ETH Zurich Feeds 100 Petabytes of NASA Data Into a Supercomputer to Speed Up Disaster Forecasting
  2. AI-Guided Portable Ultrasound Device for Battlefield Medicine Wins Tech Transfer Award
  3. AI Tool ProteinTalks Predicts Personalized Breast Cancer Drug Responses
  4. AI Scans 400,000 Reddit Posts to Find Hidden Side Effects of Ozempic-Type Drugs
  5. AI Tool Cuts Weeks of 3D Cell-Membrane Mapping Down to a Few Hours
  6. Scientists Find a Hidden Layer of Alzheimer's in How the Genome Folds
  7. Why Children Still Out-Learn AI at Language, and What It Could Teach Machines
#AI#language models#cognitive science#language acquisition
Rate this story

Related stories