Skip to content

AI in Education: Cheating, Tutoring and Assessment

4 min read · updated August 3, 2026

Institutions tend to treat this as one problem — students are cheating, find the cheaters. It is three problems with different evidence bases, and the framing determines whether the response can possibly work.

Three problems, routinely merged

  • Academic integrity. A rule was broken. This is a disciplinary and evidentiary question.
  • Assessment validity. The grade no longer supports the inference that the student can do the thing. This is a measurement question and it holds even if nobody broke a rule.
  • Learning. Does using these tools help or hinder the acquisition of skill? An empirical question, largely open.

They come apart immediately. If a course permits model use, integrity is untouched and validity may still be destroyed. If a student uses a model within the rules and learns less, that is a learning problem with no integrity dimension at all. Detection tooling addresses only the first, badly.

The merge has a predictable institutional cause. Integrity is the only one of the three with an existing owner, an existing process and an existing budget line, so a problem that is mostly about validity gets routed to the misconduct office because that is the office that exists. The result is a great deal of adjudication and very little curriculum change, which is the wrong ratio if the analysis below is right.

Integrity is really a validity problem

The concept to borrow from measurement theory: an assessment is valid to the extent that performance on it supports the inference you want to draw. An essay was never intrinsically valuable; it was a proxy, defensible because producing a good one reliably required the competencies being assessed.

That inference is what broke. If the artefact can now be produced without the competence, the proxy no longer licenses the conclusion — and this is true regardless of enforcement, because validity is a property of the inference and not of the student’s honesty. It also explains why detection cannot repair it: even a perfect detector would only tell you who broke a rule, while the grades of everyone else would still mean less than they used to, since the population of honestly produced work now includes work produced with permitted assistance of unknown extent.

Framed this way the response is not enforcement but redesign, and the question becomes which redesign restores which inference.

The redesign options and what each costs

OptionDescription
supervised conditionsRestore the inference by controlling the conditions: invigilated writing, oral examination, practical demonstration. Removes the constraint completely and is the only option that does. Costs: staff time that scales with cohort size, accessibility accommodations, and a narrowing of what can be assessed to what fits in a room and a clock.
process evidenceAssess the trajectory rather than the artefact: drafts, version history, annotated bibliographies, supervision meetings. Restores a weaker inference — that the student engaged — at moderate cost. Fabricable with effort, and the effort required is itself a deterrent.
shift the constructAssess something the tool does not do for you: critique of model output, identification of its errors, defence of a choice under questioning, application to a context only the student has. Preserves scale. Costs a genuine rewrite of the learning outcomes, and the new construct may not be the one the qualification promises.
permit and raise the barAllow assistance, require disclosure of how it was used, and set the standard where assistance is assumed. Honest and administratively simple. Costs comparability with prior cohorts, and it requires teaching staff to know what the assisted baseline actually is — which changes every few months.
two-lane assessmentA small supervised component that gates a larger unsupervised one, so the coursework grade is only credited if the supervised performance is consistent with it. Concentrates the expensive control where it does the most work. Costs design effort and produces hard cases at the boundary.

Which of these is best is not something the evidence settles, and claims that one has been shown to work should be read carefully: most published evaluations are single-institution, short, and confounded with the enthusiasm of the person who redesigned the course. What can be said firmly is the structural point — only the first fully restores the original inference, and everything else trades scale against strength.

Tutoring: what is plausible and what is over-claimed

The most-cited support for AI tutoring is Bloom’s two-sigma finding — that individually tutored students outperformed conventionally taught ones by around two standard deviations. It is worth knowing what that study actually was: small, from the 1980s, with tutoring bundled together with mastery learning, and the effect size has not been reproduced at that magnitude in subsequent work. Citing it as an available benchmark for software is over-claiming, and it happens constantly.

What is plausible on general grounds: availability at hours and volumes human tutoring cannot reach; patience with repetition; and immediate feedback, which the feedback literature does support as valuable when it is specific. What is unresolved: whether gains persist, whether model errors in a tutoring context are caught by a learner who by definition cannot evaluate them, and whether frictionless help displaces the productive struggle that the desirable-difficulties literature suggests is where durable learning comes from. That last concern is theoretically well grounded and empirically open, and it applies to a tutor that is too helpful regardless of whether it is a person or a model.

A design implication follows from that, and it is the one thing in this area with a reasonably firm basis. The tutoring behaviour supported by the learning literature is not answering; it is questioning, hinting, requiring retrieval before providing, and spacing practice. A system configured to produce complete answers on request is optimising for satisfaction, which is easy to measure, rather than for retention, which is not. Any evaluation of an educational deployment that reports engagement and self-reported helpfulness without a delayed assessment is measuring the wrong thing, and most published evaluations do exactly that.

Equity and the cost of being wrong

Two asymmetries deserve explicit weight in any policy. Detection-based enforcement produces false accusations, and published work indicates those fall disproportionately on non-native English writers — a separate page in this cluster covers the mechanism. An unfounded integrity accusation is a serious harm to a student with limited ability to rebut a proprietary score. Meanwhile supervised assessment, the only fully valid option, has its own distributional effects on students with disabilities, caring responsibilities or unstable circumstances.

Neither observation resolves the policy. Both should be on the table explicitly rather than discovered afterwards, and an institution that adopts detection without deciding in advance what evidentiary weight a score may carry has made a decision by default.

A last observation that applies to all three problems. Whatever an institution decides, students respond to what is assessed, and a rule that is announced but not reflected in the assessment design will be read as advisory. The corollary is that the cheapest honest policy is usually per-assessment rather than institution-wide: state on each task what assistance is permitted and how the work will be judged, because a blanket ban on a tool that the assessment cannot detect teaches students mainly that the rules are decorative.

AI in Education: Cheating, Tutoring and Assessment · Multigrid