News Details

img

AI Assessment Redesign

 

Classroom practices that encourage intelligent AI behaviours

The numbers no longer leave room for doubt. The 2026 HEPI survey found that 94% of undergraduates in the United Kingdom now use generative AI in assessed work, up from 3% just two years earlier.

Turnitin reports that more than half of Australian tertiary submissions between October 2025 and April 2026 contained some form of AI-assisted writing, with one in 10 more than 80% AI-generated.

Whatever assessment policies say, the unsupervised take-home essay as a test of unaided student writing is gone.

The institutional response has been retreating. The universities of Melbourne and Toronto are phasing out unsupervised take-home essays. The University of Cape Town’s AI working group wants 40% of assessments invigilated by 2027. At Durham University, a professor of philosophy stood down as chair of his board of examiners, citing a marking crisis.

But retreat is only one of three institutional responses now visible worldwide, and comparing them is instructive.

The second response, arguably the least defensible, is detection-led enforcement: running submissions through AI detectors and prosecuting the flagged.

This approach is quietly collapsing.

Detection tools produce both false negatives and, more damagingly, false positives that disproportionately affect non-native English speakers, and several institutions have faced appeals from students wrongly accused on the strength of a detector score alone.

Notably, even universities that retain detectors are stepping back from relying on them: the University of Sydney, for instance, uses Turnitin’s AI indicator only alongside other evidence, never as a standalone basis for an integrity case, on the frank institutional admission that AI use in unsupervised assessment cannot reliably be detected at all.

Good practice emerging from redesign

The third response is redesign, and it is here that genuinely transferable good practice is emerging.

The most influential model is the ‘two-lane’ approach developed by Danny Liu and Adam Bridgeman at the University of Sydney and now embedded in that university’s coursework policy. It has also been adopted or adapted by institutions from Auckland to the Australian Catholic University to VU Amsterdam, and reflected in the assessment-reform guidance of Australian regulator TEQSA, the Tertiary Education Quality and Standards Agency.

The logic is clean. Lane one consists of secure, supervised assessments; in-person written and practical tasks; oral defences; and question-and-answer sessions following presentations, all designed to ensure that each graduate possesses the knowledge and skills the degree certifies.

Lane two consists of open, unsupervised assessments in which AI use is assumed rather than policed. Since January 2025, Sydney coordinators cannot prohibit AI in unsecured assessments, because an unenforceable prohibition teaches students only that rules are theatre.

Lane two’s job is different: to drive learning and to develop and evaluate the capability students will need, working with AI critically, transparently, and well.

Coupled with programme-level design, the approach can even reduce burden. Brunel University’s programme-level assessment reforms cut the summative assessment load by roughly two-thirds.

The anxiety behind wholesale retreat by many universities is understandable. But the stampede back to the exam hall risks sidelining an important question: is what assessments were measuring actually worth measuring?

The question is how

Artificial intelligence is now part of everyday professional life. Business leaders rely on AI to analyse reports, draft documents, summarise meetings, generate ideas and support decisions.

Today’s students will graduate into workplaces where using AI is as unremarkable as using spreadsheets. A university that outright prohibits AI is preparing graduates for a world that no longer exists, and, as adoption figures show, prohibition doesn’t work anyway.

The meaningful question is whether students know how to use AI intelligently. Can they question its answers? Identify its mistakes? Recognise its biases? Decide when it is useful and when it should be ignored?

These are precisely the capabilities their future employers will demand. They are also, it turns out, capabilities we can observe, teach and assess.

Assess the conversation, not just the essay

Instead of banning AI, we should deliberately integrate it into essay assignments: encourage students to use it transparently, record their prompts, and explain how they evaluated the responses it gave them.

Their work is then assessed not only on the quality of the essay but also on how thoughtfully they used the technology.

My observation of students working this way points to a consistent pattern. Those who ask AI for answers generally produce weaker work.

Those who treat AI as a discussion partner, questioning its responses, checking facts, refining arguments and developing their own interpretations, produce significantly stronger analytical essays. The difference is not whether, but how, AI is used.

The strongest students do far more than write clever prompts. Much of the current discussion fixates on ‘prompt engineering’, but prompting alone is not what separates performance.

The best students repeatedly challenge the AI, ask follow-up questions, compare its answers against academic sources, and frequently reject its suggestions.

The AI becomes part of an ongoing conversation rather than a shortcut to submission. The weakest students accept the first response without questioning its accuracy or reasoning.

How these behaviours are actually taught

None of this happens spontaneously; the behaviours must be designed into the assignment, and in my experience five classroom practices do most of the work.

First, the AI appendix. Students submit, alongside the essay, a record of their significant AI interactions: the prompts they used, the responses they received and, critically, a short commentary on what they accepted, what they rejected, and why.

The appendix converts invisible AI use into assessable evidence of thinking. It also changes behaviour on its own: a student who knows the dialogue will be read aloud conducts it better.

Second, structured error-hunting. Early in the module, before the assessed essay, students are given an AI-generated answer to a question in the discipline and asked to find what is wrong with it: the invented reference, the misattributed theory, and the plausible but false claim.

Nothing teaches the verification imperative faster than catching a confident machine in a fabrication. Students who have done this exercise once never submit an AI citation unchecked again.

Third, the challenge requirement. The assignment brief explicitly requires students to document at least two moments in which they pushed back on the AI, questioned its reasoning, tested it against a source, asked it to argue the opposite case, and explained what the exchange changed in their thinking.

This rewards precisely the executive habit the workplace will demand: never accepting the first answer.

Fourth, a short oral defence. Ten minutes of questions on the submitted essay reliably reveals whether the thinking belongs to the student.

Those who genuinely wrestled with the material answer fluently; those who outsourced it cannot defend their own paragraphs. Oral examination scales imperfectly, but even applying it to a sample of submissions changes how every student approaches the task.

Fifth, and underpinning all of it, the marking rubric itself is rewritten so that a meaningful share of the grade attaches to the quality of the AI interaction and the verification work, not solely to the polish of the final prose. Students optimise for what is rewarded.

If the rubric rewards only the product, they will delegate the product. If it rewards the process, they will engage in the process.

Management education has taught this lesson for decades in another form: successful executives rarely accept the first report they receive.

They probe, seek evidence, challenge assumptions and consider alternatives. AI should be approached the same way, and assessment can be designed to reward exactly that behaviour.

Critical thinking just became more important

There is an irony in the AI panic: the technology that supposedly makes critical thinking obsolete in fact makes it indispensable.

Large language models invent references, produce inaccurate facts, and express incorrect information with remarkable confidence. Without careful evaluation, convincing-sounding answers can mislead uncritical students.

Stronger students naturally develop verification habits, comparing AI responses with reliable sources, identifying factual errors, and modifying AI-generated ideas before using them. Weaker students accept outputs uncritically.

That gap, not access to the technology, is what assessment should now be measuring, because it is the gap that will matter for the rest of these students’ working lives. In business, government and public policy, the value of leaders will rest less on producing information than on judging the quality of information produced by intelligent systems.

What the future essay should reward

None of this means the essay should disappear or that the retreat to invigilation is entirely wrong. In-person assessment has a place, particularly for foundational knowledge.

But abandoning the essay altogether would discard the one assessment form best suited to examining how students think over time, with sources, under realistic conditions.

Instead of asking students to summarise theories AI reproduces instantly, assignments should require them to analyse, evaluate and apply ideas to unfamiliar situations and to show their working with the machine. Assessment should increasingly reward:

• Independent judgement: positions the student has formed, defended and revised, not positions retrieved.

• Critical evaluation of AI output: documented moments where the student caught an error, challenged a claim, or rejected a suggestion.

• Evidence-based reasoning: claims triangulated against sources the student selected and justified.

• Personal reflection: the form of writing AI demonstrably cannot do on a student’s behalf.

• Transparency about AI use: students should not hide their AI interactions; they should explain them. The process record may become as important as the product.

Encouragingly, the policy landscape is already moving this way. The dominant 2026 trend across leading universities is disclosure over prohibition, and its practical forms are worth spelling out.

Tiered-use frameworks, such as the AI Assessment Scale now used in various adaptations internationally, specify for each assessment a level of permitted AI involvement, from none, through AI-assisted ideation or editing, to full collaboration, so that expectations are explicit rather than assumed.

AI-use statements, now appended to submissions at a growing number of institutions, ask students to declare which tools they used and for what.

‘AI reflections’ go a step further, asking students to write briefly about how the technology shaped their work. And at the policy level, the University of Sydney’s decision to adopt the use of AI in all unsupervised assessments, rather than maintain unenforceable bans, represents the honest endpoint of the disclosure logic.

What is still largely missing, at nearly every institution, is the final step this article argues for: assessing the quality of the disclosed interaction, rather than merely recording that it occurred.

A new definition of academic excellence

For years, many assessments rewarded information retrieval, memorisation and polished prose. AI now performs those tasks exceptionally well, which forces the deeper question universities have long deferred: what should we measure?

If higher education exists to develop judgement, ethical reasoning, creativity, self-awareness and critical thinking, then AI is not replacing education. It is helping to redefine it by making the difference between knowing and thinking impossible to ignore.

The strongest essays of the future will not demonstrate what students know. They will reveal how they think, how they evaluate evidence, how they challenge assumptions, and how wisely they work with intelligent machines.

The essay that merely proved a student could assemble information and polish prose is dead, and universities should not mourn it. It was never measuring what mattered most. The institutions that grasp this will spend the next decade designing assessments worthy of their graduates’ futures. The rest will spend it building bigger exam halls.

Nada Kakabadse is professor of policy, governance and ethics at Henley Business School at the University of Reading in the United Kingdom.

This article is a commentary. Commentary articles are the opinion of the author and do not necessarily reflect the views of 
University World News.

  • SOCIAL SHARE :