A pupil completes an excellent piece of homework. It is fluent, well structured and considerably better than anything they have produced in class. What has the assessment actually told us?
Until recently, we might reasonably have assumed that the quality of a submitted piece of work was at least a rough proxy for a pupil's knowledge and understanding. Generative AI makes that assumption much harder to sustain.
But AI is not the only reason we need to rethink assessment. Decades of research into the testing effect show that assessment can do more than measure learning. Used well, the act of being assessed can actually produce learning.
Assessment increasingly needs to do two jobs: help pupils learn and provide trustworthy evidence of what they can genuinely do.
Testing is not just measurement
The word testing often carries negative connotations in education. We associate it with grades, examinations and accountability. Cognitive science offers another way of looking at it.
The testing effect, usually described in education as retrieval practice, is the finding that trying to retrieve previously learned information can improve our ability to remember it later. Asking a pupil a question therefore does not simply reveal whether they know the answer. The act of trying to answer can change what they know in the future.
This is one of the better-established findings in learning research. A major classroom meta-analysis by Yang and colleagues synthesised 222 studies involving 48,478 learners and found a meaningful overall benefit from testing on subsequent learning.1
That gives assessment a very different purpose. A five-question quiz at the start of a lesson is not necessarily preparation for learning. It can be part of the learning itself.
Multiple-choice questions deserve another look
One of the more useful findings for teachers concerns multiple-choice questions.
It seems intuitive that short-answer questions should always produce more learning because pupils must generate an answer rather than recognise it from a list. The evidence is more interesting. Research in middle and high-school classes found that multiple-choice and short-answer quizzes both improved later examination performance when correct-answer feedback was provided.2
That does not mean question design is irrelevant. Poor distractors encourage guessing. Plausible but incorrect alternatives can also introduce misinformation, which is one reason feedback matters.3
A well-designed MCQ, however, can require pupils to discriminate between closely related ideas and can reveal specific misconceptions while allowing a teacher to sample a broad section of the curriculum quickly.
The useful question is not simply "Who chose C?"
It is: Why did eight pupils choose B, and what does that tell me about their mental model? This is where carefully designed distractors become diagnostically valuable.
For teachers, that matters. High-quality MCQs can create frequent retrieval opportunities and rich diagnostic information without generating an unsustainable marking burden.
Low stakes may be enough
There is another finding that deserves more attention. Making retrieval more consequential does not appear to make the underlying learning effect stronger.
In the large classroom review, the evidence did not show a reliable advantage for higher-stakes testing over lower-stakes testing in strengthening later learning.1 Marks, grades and consequences may still affect preparation, motivation and behaviour, but those are different questions from whether retrieving information strengthens memory.
If our objective is learning, we may not need to attach marks to every assessment. This supports a model of frequent, low-stakes assessment where getting something wrong is useful information rather than a failure.
There is an important caveat. A quiz that teachers call "low stakes" may not feel low stakes to pupils if every score is recorded, publicly compared or followed by sanctions. The design of the assessment matters, but so does the culture around it.
Feedback turns testing into learning
Retrieval is particularly useful when pupils subsequently discover whether their answer was correct. Feedback is especially important after an incorrect response, an uncertain answer or a multiple-choice question where a plausible distractor has been selected.
This suggests that the familiar sequence of question → answer → score → next question is incomplete.
If a pupil believes that acceleration means "moving quickly", simply marking an answer wrong has told them they have failed. Identifying the misconception, correcting it and asking them to apply the idea again has created an opportunity to learn.
And then we need to come back to it
Learning fades. Research on spacing suggests that retrieval is more effective when opportunities are distributed across time rather than concentrated into a single lesson or revision session.4
There does not appear to be one magical schedule that every school should follow. The useful principle is simpler: bring important knowledge back after pupils have started to forget it.
If pupils learn forces in October, perform well on an end-of-topic test and then do not encounter the material again until GCSE revision, we should not be surprised if much of it has disappeared. Assessment should therefore become increasingly cumulative, particularly for foundational, high-value and easily confused knowledge.
Remembering is not the same as being able to use knowledge
There is a danger once schools discover retrieval practice: every lesson can become a sequence of recall questions. That would misunderstand the evidence.
Retrieval is particularly good at making knowledge accessible. But being able to retrieve something does not guarantee that a pupil can use it flexibly. A pupil might know F = ma and still be unable to solve a mechanics problem because they cannot identify the resultant force. A pupil may remember historical events but struggle to build an argument about why they mattered.
Research on transfer suggests that retrieval can support application, but the benefit weakens as the final task becomes less similar to the task pupils practised.5 If we want pupils to compare, explain, evaluate, calculate or solve unfamiliar problems, they need to practise those forms of thinking too.
This is where assessment connects with deliberate practice. Retrieval helps answer: Can the pupil access the knowledge? Deliberate practice asks: Which part of their performance needs improving next?6
A useful cycle is therefore: retrieve → diagnose → practise deliberately → receive feedback → improve → retrieve again later.
Then AI changes the problem
Generative AI introduces something fundamentally different. We can now have a situation where the quality of the product increases while the amount of learning involved in producing it decreases.
A pupil can submit a beautifully structured essay without having wrestled with the argument. They can produce working computer code without understanding why it works. They can generate an explanation of electromagnetic induction that is more sophisticated than one they could independently articulate.
The product may be excellent. The evidence of learning may be poor.
The key assessment question is no longer simply "Is this good work?" It is "What can I legitimately infer about this pupil from the work they have submitted?"
This makes generative AI partly an assessment validity problem, not simply a plagiarism or behaviour problem.
We may need two lanes of assessment
One useful response is the University of Sydney's two-lane approach. It separates assessments that need secure evidence of independent attainment from assessments that are deliberately open to relevant tools, including AI.7
Trying to make an open assessment secure merely by writing "AI must not be used" at the top of the task may provide reassurance without actually protecting validity. Instead, teachers need to decide what AI use is appropriate and design the task around that reality.
Not every assessment needs the same AI rules
The AI Assessment Scale (AIAS) provides one way of making those expectations explicit. Its five levels range from controlled no-AI assessment through planning and collaboration to tasks where substantial AI use is intentionally part of what is being assessed.8
Crucially, the levels are not a ladder from bad to good. A no-AI task may be exactly what is needed if we want to know whether a pupil independently understands Newton's laws. A task involving substantial AI use may be appropriate if the outcome is the ability to critique, verify and improve AI-generated scientific explanations.
For regulated GCSE, AS and A-level coursework and non-examination assessment, school-designed AI frameworks must still sit within current awarding-body and JCQ requirements. JCQ guidance continues to require that assessed work is the candidate's own and that permitted AI-generated material is appropriately acknowledged and retained for authentication where relevant.9
We may have focused too much on the final product
AI also challenges our reliance on product-focused assessment. For an extended essay or project, perhaps the final submission should no longer carry all of the evidential weight.
For significant pieces of work, such as an EPQ, extended investigation or major essay, useful evidence might include initial ideas, planning, source selection, selected drafts, feedback, revisions, declared AI interactions and short oral discussions about key decisions.
The aim is not to create another mountain of teacher marking. It is to identify the assessments where authorship, reasoning and development genuinely matter, and where the final product alone is no longer sufficient evidence.
AI overuse may not always be primarily an ethics problem
There is also a useful executive-function lens. Sometimes inappropriate AI use will clearly be deliberate cheating. But that may not explain every case.
Consider the pupil facing several deadlines who has left everything too late, the pupil who struggles to organise a complex task, or the highly conscientious pupil whose perfectionism makes completion feel impossible. AI offers all of them an immediate route from discomfort to a finished product.
This should not remove responsibility, but it suggests that responsible AI education needs more than rules and sanctions. Pupils also need explicit teaching about planning, workload, checking AI output and recognising when AI is supporting their thinking and when it is replacing the thinking the task was designed to develop.
A useful principle is: learning first, AI later.
So what might an assessment system look like now?
The answer is not to abolish traditional assessment. AI may actually make some traditional assessment more important because schools still need trustworthy evidence of independent knowledge and performance.
But traditional assessment probably needs to sit within a broader architecture.
What this could mean for departments
- Use frequent retrieval to strengthen core knowledge and expose misconceptions.
- Make low-stakes assessment cumulative rather than leaving old content behind.
- Use secure assessment where independent competence genuinely matters.
- For substantial open tasks, collect only the process evidence that is useful and proportionate.
- When AI is permitted, assess judgement, verification and subject knowledge rather than merely the polish of the final output.
Assessment may become more important, not less
Generative AI has exposed something that was probably always true: a polished final product is not the same thing as learning.
At the same time, research into retrieval practice reminds us that assessment does not have to sit at the end of teaching as a judgement on what has happened. Assessment can be part of the mechanism through which learning happens.
A question can strengthen memory. An error can identify a misconception. Feedback can change understanding. A later question can show whether that change lasted. A secure assessment can establish what a pupil can independently do. An appropriately designed AI task can show whether they can use powerful technology without surrendering their own judgement.
What do we want this assessment to tell us, and how can the assessment itself help pupils learn?
In the age of AI, that may be a more useful starting point than simply asking how often pupils should be tested or whether AI should be allowed.
References and further reading
- Yang, C., Luo, L., Vadillo, M.A., Yu, R. and Shanks, D.R. (2021), 'Testing (quizzing) boosts classroom learning: A systematic and meta-analytic review', Psychological Bulletin, 147(4), 399–435. DOI.
- McDermott, K.B. et al. (2014), 'Both multiple-choice and short-answer quizzes enhance later exam performance in middle and high school classes', Journal of Experimental Psychology: Applied, 20(1), 3–21. DOI.
- Butler, A.C. and Roediger, H.L. III (2008), 'Feedback enhances the positive effects and reduces the negative effects of multiple-choice testing', Memory & Cognition, 36(3), 604–616. DOI.
- Latimier, A., Peyre, H. and Ramus, F. (2021), 'A meta-analytic review of the benefit of spacing out retrieval practice episodes on retention', Educational Psychology Review, 33, 959–987. DOI.
- Pan, S.C. and Rickard, T.C. (2018), 'Transfer of test-enhanced learning: Meta-analytic review and synthesis', Psychological Bulletin, 144(7), 710–756. DOI.
- Ericsson, K.A., Krampe, R.T. and Tesch-Römer, C. (1993), 'The role of deliberate practice in the acquisition of expert performance', Psychological Review, 100(3), 363–406. DOI.
- University of Sydney (2024–2026), guidance on its two-lane approach to assessment: secure in-person assessment and appropriately designed open assessment.
- AI Assessment Scale, version 2.1, official framework and implementation guidance.
- Joint Council for Qualifications, AI Use in Assessments: Your role in protecting the integrity of qualifications and current coursework/NEA guidance.
Further evidence: Gonçalves, A.O., Muniz, B.F.B. and Jaeger, A. (2025), 'Retrieval Practice Versus Elaborative Encoding: A Systematic and Meta-analytic Review', Educational Psychology Review. DOI.