What human reading reveals about AI prediction capabilities


Close-up of eye looking at text with a tech background

Artificial intelligence has become remarkably good at predicting what comes next in language. That ability helps power systems that can complete sentences, draft emails, summarize documents, answer questions and generate fluent responses from just a few words of instruction. But prediction is also part of how people understand language, which raises an important fundamental question at the intersection of AI and cognitive science: Do today's language models process language the way humans do?

Recent U.S. National Science Foundation-supported research examined that question — the answer: only partly.

Researchers found that large language models (LLMs) were good at predicting how people respond to an unexpected word while reading. But they were much less successful at explaining what happens when readers become confused and go back and reread part of a sentence.

To study that process, researchers analyzed how people read syntactically challenging sentences while high-speed eye trackers recorded where they looked and how long they spent on individual words. When reading proceeded smoothly, their eyes generally continued forward. However, when the readers encountered difficulty, they spent more time on a word or moved backward to reread something they had already seen.

The researchers were especially interested in so-called "garden path" sentences, which temporarily encourage one interpretation before forcing the reader to adopt another. For example, consider the beginning of the sentence, "The hiker found the dog …" Most readers are likely to assume that the hiker discovered the dog. But if the sentence continues, "The hiker found the dog was a delightful companion," the word "was" forces the reader to reconsider that interpretation. The intended meaning is closer to saying the hiker found that the dog was a delightful companion. The team then compared human reading behavior with predictions generated from multiple families of LLMs.

Where LLMs match humans, and where they do not

Researchers found that LLMs were good at predicting how long people would pause at an unexpected word when they kept reading forward. But when readers became confused enough to go back and reread, the models fell short.

In ambiguous sentences, people were about 92% more likely to move their eyes backward from the key word and 102% more likely to do so at the next word. The largest model predictions were only 8% and 21%, respectively. The difference was even larger in the time spent rereading. People spent an additional 584 milliseconds rereading after the confusing part of a sentence. The largest model prediction was just 29 milliseconds.

The findings challenge an increasingly influential idea about both human cognition and AI. LLMs can produce extraordinarily human-like language, making it tempting to conclude that their underlying computational mechanisms resemble human language processing. This research provides evidence that the similarity has limits.

What's next

Researchers are exploring how these findings could help AI assistants communicate more clearly.

For example, someone using an AI assistant to install a wall-mounted shelf might receive an instruction that seems clear to another AI system but is ambiguous to a person, such as whether to drill before or after locating a wall stud. Researchers call this kind of testing "user simulation," in which one AI system models how a person might respond to another AI system's instructions. But if AI systems do not get confused in the same places people do, those simulations may miss problems human users would encounter.

A better understanding of the difference between AI prediction and human comprehension could help researchers build systems that recognize when instructions may be unclear rather than simply continuing with a plausible response. That could improve AI assistants, tutoring systems and other tools that need to communicate complex information clearly.

This NSF-supported work is an example of foundational research that can help lay the scientific groundwork for AI technologies that are more reliable, useful and effective in Americans' everyday lives.

The project was partly supported by NSF Directorate for Computer and Information Science and Engineering awards IIS-2504953 and IIS-2504954