News

Bringing AI into test development — with foundations first

  • Faculty of Humanities, Education and Social Sciences (FHSE)
    Luxembourg Centre for Educational Testing
    21 April 2026
  • Category
    Research
  • Topic
    Computer Science & ICT, Education & Social Work

Over the past years, LUCET has developed cognitive models that help explain how students solve assessment items in the annual school-monitoring programme , making item difficulty more predictable and potential cultural or linguistic bias easier to spot at the design stage. That work also laid the groundwork for what comes next: bringing AI into the ÉpStan test-development process in a way that holds up scientifically. 

That order matters. AI promises speed and scale, but on its own those gains stay shallow and can undercut the quality criteria that make an assessment trustworthy. With a cognitive-model foundation in place for parts of our item development, we can now start using AI operationally as a tool that augments our test developers rather than replacing their judgement. 

Preparing for the European AI Act 

Educational assessment is classified as high-risk under the European Union’s AI Act. To meet that standard, LUCET has built a governance framework and published a white paper, , offering operational guidance for the responsible use of generative AI in large-scale assessment. It turns abstract regulatory principles, such as human oversight, transparency, fairness, data protection, and AI literacy, into concrete commitments tied to the psychometric quality-assurance practices we already follow. An internal LUCET strategy then translates this into day-to-day rules for the ÉpStan test-development workflow with special consideration of avoiding AI-based copyright infringement. Together they put Luxembourg among the early movers in Europe working to make the AI Act practicable for education, rather than merely compliant on paper. 

AI as a principled extension, not a shortcut 

These cognitive foundations are also what make AI integration scientifically defensible. In a study published in , LUCET researchers tested whether GPT-4 could produce German reading-comprehension texts good enough for national assessment. Crucially, the texts weren’t generated freely: they were built from a Text Analysis Cognitive Model (TACM) â€” a structured template, defined by a subject-matter expert, that fixes the cognitive and linguistic features a text must have. 

Text analysis cognitive model (TACM). The text features are based on the requirements defined by expert test developers and subject experts for Luxembourgish 5th grade

In a blind review, 89 readers rated human- and AI-written passages on readability, correctness, coherence, engagement, and fitness for assessment without being told which was which, and could not reliably tell them apart. But the results were not uniform: giving the model an example text to work from yielded strong informative passages, while human authors kept the edge for narrative texts, where coherence and engagement matter most.

A parallel SWOT analysis with our own test developers named the benefits, such as efficient first drafts, adapting content across difficulty levels, easing multilingual demands, alongside the risks that keep a human in the loop: factual errors, cultural and pragmatic bias, and the pull to trust a fluent-sounding text without scrutinising it. 

That is the approach in one word: augmentation and oversight. AI extends expert judgement and draws its quality from the cognitive foundations underneath it. The research continues, and operational use in the ÉpStan workflow is now beginning. 

Share this