Academic assessment

When AI can produce the assignment: how to redesign assessment without designing a surveillance regime

A humane, practical approach to authentic assessment that values process, judgement and dialogue instead of escalating detection and monitoring.

← All articles

The first reaction to generative AI in education was understandable: protect the assignment. Institutions added declarations, detection tools and more monitoring because a familiar piece of coursework could now be produced in seconds. Yet the deeper problem was never the document. It was that educators had been using the finished document as a convenient proxy for learning, judgement and authorship.

UNESCos recent discussion of assessment in the AI age frames the disruption as an opportunity to reconsider what is worth measuring. That is the productive direction. The goal is not to make every task AI-proof; no stable design can promise that. It is to collect better evidence of what a learner understands, how they make decisions and whether they can use tools responsibly.

Decide what remains distinctively the learners work

Start with the learning outcome, then identify the human performance you need to see. It may be selecting a defensible approach, interpreting conflicting evidence, noticing an ethical problem, explaining a trade-off, adapting to feedback or applying knowledge in an unfamiliar context. Those are stronger anchors than write 2,000 words.

For each task, define which assistance is permitted. AI might be allowed for brainstorming but not evidence selection; for language editing but not analysis; or throughout the task if the learner must critique the output and disclose important decisions. Avoid a single institution-wide rule that treats every discipline and outcome alike. A tool appropriate in a software engineering exercise may undermine a task designed to assess unaided language production.

Explain the rule with examples. Use AI responsibly is not operational guidance. Show what acceptable acknowledgement looks like, what records to keep, how generated claims should be checked and where assistance would prevent the outcome from being assessed.

Replace the one-shot artefact with an evidence sequence

A finished essay gives the marker one surface to inspect. A sequence creates a richer and more teachable record. Ask for an early problem framing, a source choice with rationale, a plan, selected working notes, feedback response and final product. Not every stage needs a mark. The purpose is to make important decisions visible.

Version history can help, but it should not become forensic surveillance. Tell learners what will be collected and why. Focus on meaningful changes rather than keystroke patterns. A short reflection can explain where the approach changed, what feedback mattered and what the learner would do differently. The marker can then compare that account with the artefacts.

This approach also supports learning. Feedback arrives while there is time to act, and educators can see misconceptions before they harden into a polished final answer. If generative AI contributed, the learner can demonstrate the judgement used to accept, reject or revise it.

Use dialogue as verification and learning

A five-minute conversation can reveal more than a detector score. Ask the learner to explain one decision, apply the idea to a new example or respond to a reasonable challenge. The aim is not a hostile viva designed to catch them out. It is a proportionate check that the submitted work connects to understanding.

Use a short question set and scoring guide so the conversation is consistent. Offer an accessible alternative where live speech creates a barrier. Record the outcome and rationale, not an unnecessary video archive. Dialogue can be sampled for low-consequence work and used more consistently where authorship or competence is central to the outcome.

In large classes, verification can be targeted. Random sampling discourages outsourcing without implying suspicion. A staged task can route only inconsistent or incomplete evidence to a conversation. Make the routing rule transparent and give learners a fair opportunity to respond.

Assess AI use when AI use belongs in the discipline

In many fields, graduates will work with AI. Pretending the tool does not exist can make assessment less authentic. Instead, assess the capability to use it well. Ask learners to formulate and refine a request, check sources, identify unsupported claims, compare outputs, protect sensitive information and explain where human judgement changed the result.

Do not award marks for elaborate prompting alone. The valuable evidence is in the quality of the goal, verification and decision. A learner should be able to explain why an output was appropriate, what risk remained and how they would proceed if the tool were unavailable.

Keep unaided elements where they are genuinely needed. A nurse, engineer or accountant may need core knowledge without external help in time-critical contexts. Programme-level design can combine controlled checks, authentic assisted work and dialogue rather than forcing one task to prove everything.

Do not turn detector output into a verdict

AI detectors produce a signal, not proof of misconduct. Text can be misclassified, and a probability says little about the circumstances in which the work was created. Treating a flag as a finding risks unfair accusation and can disproportionately affect writers whose style differs from the detectors training patterns.

If a tool is used, define a narrow purpose and validate it on relevant work. Keep the output away from the markers initial academic judgement where possible. A flag may prompt examination of the evidence sequence or a conversation, but the institution should follow its normal fair process and consider the learners explanation. Never ask a student to prove a negative solely because a system produced a confident number.

Monitor outcomes: how many flags are confirmed through independent evidence, which groups are affected, how much staff time is consumed and whether the tool changes trust in the classroom. Stop using it if the benefit does not justify the harm and workload.

Make academic integrity guidance teachable

Learners are receiving conflicting messages. One module encourages AI experimentation while another prohibits tools without explaining what counts as use. Create task-level guidance in the assessment brief and repeat it in the submission flow. Give examples of permitted editing, prohibited generation, collaboration and acknowledgement.

Teach source verification and disclosure before grading them. Students cannot be expected to infer a new academic convention. Show how a generated claim can be traced to a credible source, how confidential placement data must be protected and how to describe substantive assistance without pasting an unreadable transcript.

Staff need the same support. Markers should know what evidence is available, how to avoid overreacting to stylistic cues and where to escalate a concern. Programme teams need a forum to compare cases so rules do not drift between modules.

Design fair routes for different access to tools

If AI use is required, ensure all learners can access an approved tool and understand its data terms. Paid subscriptions, device performance, language support and regional availability can create unequal conditions. Provide an alternative when a learner cannot or reasonably chooses not to use a particular provider.

Accessibility deserves special attention. Generative tools may help some learners express ideas, but interfaces and outputs can create new barriers. Do not interpret use of assistive technology as suspicious behaviour. Adjustment information should remain protected, and an integrity process should consider how approved support shaped the work.

A redesign workshop that produces something useful

  1. Write the learning outcome and the decision the result will support.
  2. Identify the observable human judgements or performances needed.
  3. List how AI could support, distort or conceal each one.
  4. Choose a balanced evidence sequence: artefact, process, dialogue and selected controlled work.
  5. Define permitted assistance and disclosure with examples.
  6. Create scoring criteria for quality, reasoning and verification rather than surface polish.
  7. Test workload, accessibility and fairness with representative learners and markers.
  8. Review the design after delivery using feedback, concerns, appeals and evidence quality.

Do this at programme level where possible. If every module independently adds an oral check, process portfolio and declaration, the cumulative burden will be unmanageable. Map where each capability is best evidenced and let tasks share the work.

What a better assessment feels like

A learner should know what they are meant to demonstrate, where tools may help and what remains their responsibility. A marker should receive more than a polished artefact: they should see enough process and explanation to make a reasoned judgement. An integrity concern should open a fair inquiry, not trigger an automated verdict.

Generative AI has made weak proxies easier to see. That is uncomfortable, but useful. Assessment can move away from policing the production of a document and toward evidence of thought, application and responsible tool use. The result is not only more resilient to AI. It is more closely connected to learning.

Sources and further reading

Primary guidance used for the current facts in this article. Always confirm requirements for your jurisdiction and use case.

Topic FAQ

Questions about academic assessment

Should students be prohibited from using generative AI?

A blanket rule is rarely sufficient. Define permitted and prohibited assistance for each task, explain why, and design evidence that still reveals the learner's judgement.

Can AI detectors prove academic misconduct?

Detector output should not be treated as proof on its own. Human inquiry, assessment evidence and fair procedures remain essential.

What is an authentic assessment?

It asks learners to apply knowledge through decisions, creation, explanation or reflection in a context connected to the intended learning outcome.

How can educators verify authorship without invasive monitoring?

Use staged artefacts, short conversations, version history, personalised application and comparison with in-class work.

Ready when you are

Turn assessment evidence into a decision you can explain.

See how the platform connects design, delivery, evaluation, publication and capability reporting.

Book a demo View sample report