Skip to content

Designing Assignments in a World With AI

10 min read · updated August 4, 2026

You cannot detect your way out of this, and the attempt does real harm. What works is designing tasks where the shortcut produces a visibly worse answer — which is a solvable design problem, and mostly a question of where the task gets its inputs.

Detection is the wrong goal

The case against detection is not squeamishness. Detectors are wrong in both directions, their errors fall disproportionately on non-native speakers and formulaic writers, and an accusation carries a cost to a student that no false-positive rate justifies. The arithmetic of what a small error rate does across one cohort is worked out in spotting AI-written text, and the documented cases are in AI detection in schools.

There is also a design argument, which is the one that persuades colleagues. If a task can be completed to a passing standard by a machine with no access to your course, your students or your data, then the task was measuring the production of a document rather than the learning it was supposed to evidence. That was true before this technology existed; it was merely harder to notice.

Take the position openly. A blanket ban that cannot be enforced teaches students that rules are theatre. A considered rule about which uses are permitted, enforced by task design rather than by suspicion, is defensible and survives contact with a cohort.

Four properties of a resistant task

The question to ask of any assignment: what does a person need that a model does not have? There are four durable answers, and a task with two of them is already hard to outsource usefully.

PropertyDescription
local inputThe task requires material that exists only in your setting: this week's seminar discussion, the data the class collected, a specific artefact in the room, the version of the case study you handed out with three numbers changed.
personal stakeThe answer depends on the student's own decisions, errors or experience. 'Analyse the mistake you made in the lab on Tuesday and what you would change' has exactly one author who can write it.
defensible in personThe submission is followed by a short conversation, a presentation, or an in-class extension. The artefact stops being the assessment and becomes the material for it.
visible processDrafts, annotated sources, a research log or version history are part of what is marked. Not to catch anybody - because the process is where most of the learning is.

Note what is absent from the list: difficulty. Making a question harder does not help, because the shortcut scales with difficulty at least as well as the student does. Obscurity does not help either. Specificity to your context is the property that holds.

Three assignments, redesigned

Essay: causes of an historical event

Before. “In 1,500 words, discuss the causes of the 1926 general strike.” This is answerable to a solid 2:2 standard from general knowledge in under a minute.

After. “Three of the six sources in this week’s pack disagree about the role of the coal owners. In 1,200 words, establish which of the three is least well supported by the other three, quoting from the pack. Then, in 300 words, describe what additional evidence would change your conclusion.”

The pack is yours, the disagreement is one you selected, and the second part asks for a judgement about evidence that does not exist. A model given the pack still helps — and that is acceptable, because the student has to supply the pack and evaluate the output against sources they can see.

Report: analyse a dataset

Before. “Analyse the attached sales data and report your findings.”

After. “Here is the dataset. It contains three deliberate data-quality problems. Find them, document how you found each, and state what your headline figure would have been if you had not. Submit your working file with the steps visible.”

Planted faults are the most reliable trick in this whole area. The shortcut path produces a clean, confident analysis of dirty data, which is precisely the failure the assignment is now testing for. It also teaches the single most important habit in data work, and the verification method generalises — see making sense of a spreadsheet.

Reflective or practical writing

Before. “Reflect on your placement experience.” Generic reflective writing is the single easiest thing to generate.

After. “Choose one five-minute interaction from your placement that went badly. Describe it in the present tense. Identify the point at which it could have gone differently and what you would have needed to know. Bring the write-up to the seminar; you will exchange it with a partner and answer their questions on it.”

Specificity plus a person in the room. The seminar exchange costs no marking time and does more assessment work than a rubric.

Assessing the process, not only the artefact

Shifting some weight from the final document to the work that produced it is the most transferable change available, and it is worth doing on its own merits.

  • A required annotated bibliography, submitted a fortnight early. Three sources, with two sentences each on what the source argues and why it is relevant. Fabricated sources become visible here, at a stage where it is a conversation rather than a misconduct case.
  • A ten-minute checkpoint. A plan, a paragraph, a problem the student has hit. Cheap, and it makes the eventual submission predictable.
  • Version history as a submission requirement where students already write in a cloud document. State the requirement in advance and apply it to everyone, or it becomes surveillance of the suspected.
  • A short viva on a sample. Five minutes, on a randomly drawn subset, announced in advance. The deterrent comes from the possibility, not the coverage, and the honest students find it easy.

Write the rule down

Most disputes are caused by the absence of a rule rather than by its breach. Write one per assignment, in one sentence, on the brief itself — a policy in a handbook nobody opens is not a rule. A workable three-tier scheme:

AI USE ON THIS ASSIGNMENT: TIER 2

Tier 1 - none. No AI assistance of any kind. Used where the task is
  measuring unaided production (in-class writing, language
  acquisition, the exam).

Tier 2 - permitted for preparation and revision, not for producing
  submitted text. You may use it to test yourself, to check grammar,
  to explain a concept. Sentences you submit must be yours.

Tier 3 - permitted throughout, with a 200-word appendix stating what
  you used it for, one prompt you found effective, and one place its
  output was wrong. The appendix is marked.

If you are unsure whether something falls inside the tier, ask me
before you do it. Asking is never penalised.

The tier-3 appendix is worth more than it looks. Asking a student where the output was wrong requires them to have evaluated it, which is the competence you would want them to leave with anyway, and it is a reasonable thing to teach explicitly — the structure of a session for that is in teaching a group to use it well.

What this costs, honestly

These redesigns cost time. A bespoke source pack takes an afternoon to assemble and dates in three years. Planted data faults need making and documenting. Checkpoints and vivas cost contact hours that are already allocated. Anyone claiming this is free has not run it across a hundred and eighty students.

Two things reduce it. The redesign is per module rather than per year, so the cost amortises. And the marking usually gets faster: specific tasks produce specific answers, which are quicker to mark than a hundred and eighty competent generic essays that all say the same thing.

What it does not cost is the enforcement overhead of a detection regime: the appeals, the meetings, the students who disengage after being wrongly accused, and the institutional exposure when one of those findings is challenged. Set against that, an afternoon on a source pack is cheap.