What would count as evidence?
Why AlignED Studies is organised as five questions.
The delegation is already happening. Schools and platforms use AI for grading, feedback and planning advice, mostly on the strength of general capability rather than evidence from the tasks themselves. General capability turns out to be a weak guide to performance on specific professional tasks.
Report 1 measured this directly. Across five evaluations, performance on one task did not predict performance on another, and the same models that compared student work accurately reached only fair agreement with experts when grading against standards. Trust has to be earned per task, which is why this site is organised as questions rather than a leaderboard.
The five questions cover the delegation end to end. Can models judge student work? Do they hold the knowledge teachers are certified on? Does what they produce follow the evidence they can recite? How is AI actually being used in education? And can research like this itself be done with AI? Each question has its published reports and work in preparation listed on the studies page.
The method stance is constant across all of it. Pilots are labelled as pilots, data and methods are open, and each study states its main limitation up front.