Every large-scale assessment runs on a constant, huge supply of test items. Traditionally, every single one is written by a small team of experts, reviewed, and then field-tested on real students before it counts for anything – which eats up a lot of time and money. For years at LUCET we have been working on a way to build fair, high-quality maths items at scale without that overhead, using what we call cognitive item models: formal descriptions of the mental steps a task is meant to trigger.
In a large-scale study just published in , we built 48 such models and used them to generate 612 language-reduced, image-based maths items for Grades 1, 3 and 5. We also built a prototype generator for this, autoMATH (automatic math item generator), which turns one cognitive model into many equivalent versions of an item. With data from Luxembourg’s school monitoring programme (over 35,000 students), we found that most of what makes an item hard comes down to the cognitive features we deliberately built into it: number range, carrying, or whether objects are laid out horizontally or vertically, for instance.
Philipp Sonnleitner presenting the project at Luxembourg’s LuxDoc Science Slam 2025
Item model with six of the twelve generated and administered test items. The first row varies cognitive factors while keeping the context constant, whereas the second row varies the contextual embedding while maintaining identical cognitive characteristics
This changes how item development works. It becomes more scalable and predictable, and fairer already at the point of construction. Rather than uncovering bias only after a test has run, we can check whether particular item components disadvantage certain student groups while an item is still on the drawing board. With a method we call Differential Radical Functioning, we could probe fairness across cultural and linguistic backgrounds directly at the level of item design — and found that differences by cultural background can persist even when language proficiency is accounted for.
To check that the models actually reflect what children do, we validated them empirically, among other things with eye-tracking, watching where young learners direct their attention as they solve a task.
From a single score to a real learning profile
The same cognitive models do a second job: they turn test results into diagnostic feedback for teachers and learners, via Cognitive Diagnostic Models (CDMs). A large-scale assessment usually gives back just one number per subject — rarely enough to know what a child can already do and where to help next.
In early numeracy (published in ), we traced how first-graders progress through four skills — counting, addition below 10, decomposition, and addition above 10 — and found a clear order: 93% could count objects, but only 36% reliably solved additions above 10. Decomposition emerged as a pivotal threshold skill, linked to both competence and confidence. In a follow-up with TU Dortmund (), hierarchical CDMs built from the same models sorted Grade 3 students into coherent, interpretable proficiency profiles that line up with how mathematical ability tends to develop.
Example feedback for a student giving information about the relative standing within the student cohort and a more detailed individual evaluation of skills mastery including strategies to improve
The approach is now moving into German and French reading assessment in secondary schools, where the ÉpStan Pre-Level 1 pilot is designed to capture what struggling readers can do, rather than only where they fall short, giving useful feedback to the learners who need it most.
Parts of this research was funded by the FNR (grant FAIR-ITEMS C19/SC/13650128).