Recent research conducted by the University of Oxford and the Allen Institute for AI (Ai2) reveals a surprising similarity between how humans and AI systems, such as ChatGPT, process language. This new study, published in the Proceedings of the National Academy of Sciences, illustrates that large language models (LLMs) generate language through analogy rather than by strictly following grammatical rules.
Traditionally, it was assumed that LLMs derived language from the rules embedded in their training data. However, this research indicates that they rely more heavily on examples and memories, paralleling human behavior. The study spotlighted the pattern of suffixes “-ness” and “-ity,” used to transform adjectives into nouns. By presenting LLMs with artificial adjectives like “cormasive” and “friquish,” researchers assessed how these models generalize language based on analogy rather than rules.
Remarkably, GPT-J, the LLM utilized for the study, processed these novel words by drawing parallels with known words from its training—mirroring how humans might handle a similar linguistic task. This analogy-driven approach is further evidenced by the model’s predictions aligning closely with the frequency and patterns in its prior data exposure, effectively treating each new word as a memory trigger.
Despite this capability, the study highlights a key difference between LLMs and human language processing: humans organize linguistic knowledge into a coherent ‘mental dictionary,’ while LLMs process every instance separately without consolidating them into a single entry. This limitation suggests why LLMs require significantly larger datasets to achieve proficiency compared to human learners.
Senior author Janet Pierrehumbert from Oxford emphasizes that while these AI models are impressively skilled, their abstraction level doesn’t match human thinking, explaining their extensive data dependencies. Dr. Valentin Hofman from Ai2 and the University of Washington points out the synergy achieved between linguistics and AI in this study, offering clearer insights into LLM functionality that could drive future AI advancements.
In conclusion, the research sheds light on the nature of AI language generation, emphasizing the importance of analogy and stored examples over traditional rule-based processing. This paradigm provides valuable insights into AI development and its alignment with human linguistic capabilities.
Key Takeaways
- Analogical Reasoning: LLMs like ChatGPT use analogy and stored examples, rather than strict grammatical rules, to generate language.
- Comparative Understanding: LLMs show similarities to human analogical reasoning but lack a unified mental dictionary.
- Data Dependency: The study underscores LLMs’ need for bigger datasets compared to humans to achieve linguistic sophistication.
- Interdisciplinary Insights: Insights from this research bridge AI and linguistics, paving the way for more robust and explainable AI models.