Using language requires a myriad of cognitive mechanisms to work together seamlessly. Among them, a particularly intriguing phenomenon is the human capacity for syntax. The present thesis focuses on two cognitive mechanisms that are considered to play a crucial role in syntactic processing: rule abstraction and hierarchical processing. Both mechanisms have been empirically investigated through artificial grammar learning (AGL), a well-established psycholinguistic experimental paradigm for studying the learning of abstract structural rules as a precursor of natural language syntax. In recent years, Transformer-based large language models (LLMs) have emerged as the most successful artificial systems for natural language processing to date. Yet their inner workings, which are based on token prediction without explicit grammatical representations, and the limits of their capabilities remain insufficiently understood, as do the implications they might have for cognitive models of language and theories of human syntax. The present thesis addresses two research gaps: first, the focus of artificial grammar learning (AGL) research on language comprehension while overlooking language production; and second, the lack of valid comparisons between humans and large language models (LLMs) in terms of their AGL capabilities. To this end, a novel production-based experiment, which builds on prior AGL work with LLMs targeting the mechanisms of rule abstraction and hierarchical processing, is designed and carried out with adults. The results demonstrate that humans significantly outperform LLMs in a matched AGL setting; hierarchical abstraction poses a particularly challenging task for LLMs, which seem to rely on linear, rather than hierarchical structural regularities. Human performance, in turn, is stable across non-hierarchical and hierarchical grammar learning tasks, with somewhat lower success rates in the latter. The findings of the present thesis contribute both to AGL research in humans and to the growing literature on the cognitive plausibility of LLMs and human–LLM comparison.
|