Technology
Research into the AI data efficiency gap explores how children outlearn models
While large language models require billions or trillions of words to achieve fluency, human children master language on a fraction of the data.
The short version
- Human children typically achieve fluent language mastery after hearing roughly 10 million to 100 million words, whereas state-of-the-art large language models require billions or trillions of tokens.
- Researchers refer to this vast disparity in the quantity of information required to learn language as the data efficiency gap.
- Scientists are using baby-scale language models to test theories of human language acquisition and find ways to build more data-efficient artificial intelligence.
Key facts
- Human children can achieve fluent language mastery after hearing roughly 10 million to 100 million words, whereas large language models require billions or trillions of tokens.[MIT Technology Review]
- The vast difference in data requirements between children and artificial intelligence is referred to by researchers as the data efficiency gap.[MIT Technology Review]
- The BabyLM competition, launched by researchers including Alex Warstadt and Leshem Choshen, challenges participants to train language models on a child-scale corpus of 10 million to 100 million words.[MIT Technology Review]
- In the 1950s, linguist Noam Chomsky argued that children possess innate knowledge of grammar to overcome the 'poverty of the stimulus,' challenging B.F. Skinner's view that language is learned entirely through environmental conditioning.[MIT Technology Review]
What remains uncertain
- The exact biological or cognitive mechanisms that allow human babies to learn language from so little data remain a mystery.[MIT Technology Review]
Sources
- Kids outlearn AI—and we still don’t know whyMIT Technology Review metered