Studies
The work is organised as five questions. Published reports sit under the question they address, alongside work still in preparation. Each study states its status, the exact models tested, the data used, and its main limitation up front. Pilots are labelled as pilots.
Direction 01
Can models judge student work?
Schools already use AI to grade and give feedback. These studies measure how closely model judgements agree with expert judgement, and how the task format changes that agreement.
Benchmarking AI Models for Educational Practice
Five evaluations across 32 models from five providers. On judgement of student work, models compared paired work accurately, but on standards-based grading the best chance-adjusted agreement with experts was κ = 0.25. No single model ranks first on every task, and performance on one evaluation does not predict performance on another.
Read the study →Can LLMs Assess Complex Student Competencies?
Two Gemini models scored 18 student essays against rubrics matched to each corpus, with expert scores as the reference. A contamination test found both models could recall published scores for the well-known ETS corpus without seeing the rubric. Assessment accuracy on public corpora may be inflated by score memorisation, which argues for contamination testing as standard practice.
Read the study →- Grading against official curriculum achievement standards, testing whether model scores cluster at the middle of the scale and shift with the position of the work in the prompt. In preparation
- In a pilot, models identified none of the above-satisfactory work when classifying it directly, yet compared pairs of the same work with high accuracy. A follow-up tests whether comparative judgement protects LLM evaluators from surface-level quality inflation. In preparation
Direction 02
Do models hold what teachers need to know?
Professional teaching knowledge differs from the school subjects most benchmarks test. These evaluations draw on neuromyth research and teacher certification exams.
Benchmarking AI Models for Educational Practice
Report 1’s knowledge evaluations cover 32 neuromyth items adapted from Dekker et al. (2012), 12 diagnostic reasoning scenarios, and 1,143 teacher certification items spanning general pedagogy and inclusive education. On an eight-item confidence probe, no model expressed uncertainty, even when answering incorrectly.
Its judgement evaluations are listed under Direction 01.
Direction 03
What do models produce when generating freely?
A model can state the evidence on a direct question and still not apply it when generating. These studies compare what models produce with what the evidence supports.
Do LLMs Fade Worked Examples?
Six models generated worked-example sequences under three prompting conditions. All six reference fading in their reasoning traces when told to apply cognitive load theory, yet none applies it in the actual output unless the prompt spells out what fading is. Extends an earlier two-model pilot to six models, including three open-weight models.
Read the study →- Thirteen models generated teaching advice from 20 open prompts. Whether the output sits closer to evidence-based or non-evidence-based practice is measured by embedding distance to two reference corpora. Analysis underway
- LLM responses and real teacher responses to the same moments in recorded maths lessons, compared turn by turn. In preparation
Direction 04
How is AI actually being used in education?
Benchmarks test what models can do. Usage data shows what people do with them.
AI and Education: What 152,000 Conversations Reveal
A descriptive analysis of the 152,088 education-related conversations in the Anthropic Economic Index V4. Students are the primary users, with 59.5% of education usage classified as coursework. Directive requests and task iteration dominate, while feedback loops are nearly absent at 1.6% of interactions.
Read the study →Direction 05
Can this research itself be done with AI?
Much of this research is done with AI assistance, so the trust question applies to the methods as well. The tools below are built to keep their reasoning traceable.
- Aligned Interviewer AI-led qualitative interviews, synthesised against the researcher’s own framework. By invitation.
- Landscape Scanner field scans where every claim traces to a verbatim source quote. Source on GitHub
- Analytical thinking continuum an empirically derived progression of analytical thinking, and a Socratic interview built on it.
- Qual Research Suite the engine shared by the above, being consolidated into one tool. In development.
- A methods paper on deriving learning progressions from comparative judgement, with held-out testing of the resulting level descriptors. In preparation