A Fair Process for Assessors When AI Use Is Suspected
July 23, 2026 · Programmatic SEO OS
A Fair Process for Assessors When AI Use Is Suspected
By the FiftyGPT Editorial Team, reviewed against our responsible-use editorial standard. Last updated 23 July 2026. Educational guidance for assessors, not legal advice. Questions or corrections: contact us.
A fair process when AI use is suspected starts with one rule: a detection score is a weak signal, never a verdict. When a score or another cue makes you suspect AI use, pause before any accusation, preserve the student's drafts and version history, gather corroborating evidence, then hold a calm, non-accusatory conversation before you conclude anything. Human review belongs at every stage, because detectors estimate probability, not authorship, and they flag conventional and non-native English writing more often than people expect. A score alone can never be the sole basis for an academic integrity finding.
FiftyGPT is a free platform of 105+ AI writing and academic tools, built around an AI detector that estimates the probability that text was machine-generated, for students, teachers, writers and developers who want detection handled responsibly. Its stated position is plain: detector results are probability signals, not proof, to be weighed alongside writing history, citations, drafts and human review. What this guide adds, beyond a score-interpretation article or a student's "what to do if flagged" page, is a procedure an assessor can actually run: a worked timeline, a triangulation example, a weighted evidence table, an outcomes ladder and a sample case record.
What a fair process protects
A fair process protects two people at once: the honest student from a wrongful accusation, and you from a decision you cannot later defend. It does not set out to prove cheating; it reaches a conclusion that stays proportionate to a real person whose record is at stake. Hold on to one point if your first worry is accusing a student who did the work: a high score is not an accusation and it is not proof of cheating. It is a prompt to look properly, and every stage below is built so a hunch never becomes a finding on its own. Assessors carry a real duty of care, because a mistaken accusation wounds an honest student. A detector only measures how statistically predictable finished text looks; it does not see who sat at the keyboard, and it cannot.
Before you act: a short preparation checklist
Most unfair outcomes are set in the first hour, when a surprising score triggers a hasty message. Run these checks before you make contact.
- Open your institution's policy first. Follow the approved misconduct and evidence procedures where you work; this guide supports that policy, it does not replace it, and it is not legal advice.
- Name the signal honestly. A single score, a shift in a student's usual voice, or an odd submission pattern is a reason to look closer, not a finding.
- Find the process evidence early. Check whether you can reach drafts, revision history, notes or earlier authentic work before you form a view.
- Question your own assumptions. Be candid about whether this student writes English as an additional language, which is flagged disproportionately, and about any prior impression you hold.
- Decide who reviews alongside you. Line up a second colleague or integrity officer before anything escalates, so human review never rests on one stressed marker.
The five-stage process, step by step
This is a repeatable flow from a first signal to a proportionate conclusion, with human judgement inside each stage.
- Signal. Write down what triggered your concern and the date, and treat it as a question to investigate, not an answer you already hold.
- Pause and preserve. Do not accuse or hint. Keep the submission, drafts and version history exactly as they are, so nothing is lost before review.
- Gather and corroborate. Assemble the evidence set below, looking actively for signs the work is authentic, not only for signs against it.
- Conversation. Invite the student to a calm, non-accusatory talk and ask them to walk you through their process, sources and how the draft took shape.
- Proportionate conclusion with human review. Weigh the full picture with a second reviewer. The outcome may be no action, a learning conversation about disclosure, or a formal referral where corroborated evidence supports it. The score never makes this call.
For the wider context, our practical guide to AI detection for teachers shows how detection sits in an assessor's day-to-day workflow.
A worked example: reading the timeline from version history
Process evidence stays abstract until you watch it work. Suppose a 1,500-word reflective essay is flagged, and its version history shows the document created at 9:02 pm and barely touched until 11:47 pm, when a single revision pasted in roughly 1,400 words at once, followed by three tiny edits and submission at 11:58 pm. That timeline, not the score, is what earns a conversation, because authentic drafting rarely arrives as one block eleven minutes before a deadline.
Now invert it. The same essay might instead show forty saved revisions across nine days, sentences reordered, a paragraph deleted and later restored, and a word count that climbs and dips as the argument is reworked. A high score sitting on top of that trail points toward authentic writing. The shape of the editing history carries far more weight than any single number, and it can point either way.
How three weak signals combine into one defensible picture
Triangulation means several independent signals agreeing, not one number shouted louder. Treat each signal as pointing toward AI use, toward authentic work, or nowhere, and read them together.
Take the same flagged essay. Suppose the detector returns 82% (leans weakly toward AI), the version history shows the one-paste 11:47 pm timeline (leans toward AI), and two cited sources cannot be found in the library catalogue (leans toward unedited generated text). Three independent signals now point the same way, and no single one carries the decision. Change one variable and the picture changes with it: keep the 82% but replace the timeline with forty revisions over nine days and add real, correctly used citations, and two of the three signals now point toward authentic work. The number never moved; the decision did.
Corroborating evidence to weigh alongside a score
A fair conclusion rests on that convergence rather than one number amplified. The table below ranks common evidence types by how much weight each can reasonably carry on its own. Process evidence is strong because it shows how the work came to exist; a score is weak alone because it only describes the finished text.
| Evidence type | What it can show | Weight on its own |
|---|---|---|
| Drafts and revision history | Whether the work developed over time in a human way | Strong |
| Version history and document metadata | Editing timeline, large pasted blocks, unusual jumps | Strong |
| Comparison with earlier authentic work | Whether voice, level and habits are consistent | Moderate to strong |
| The student's own explanation of their process | Understanding of their sources, argument and choices | Moderate to strong |
| Citations and source trail | Whether references are real, relevant and correctly used | Moderate |
| A detection score | A probability estimate of predictability, not authorship | Weak alone |
The strongest items are all forms of process evidence. For how to read the number itself before you weigh it, see our companion piece on interpreting a detection score responsibly, and the Plagiarism Risk Checklist helps you separate genuine misconduct from over-close paraphrasing or a missing citation, a different problem that calls for a different response.
Matching the outcome to the evidence
A proportionate outcome is the one the evidence can actually carry. Use this ladder to match a realistic evidence picture to a defensible response, then confirm the choice with your second reviewer.
- Score high, process evidence supports authorship. Rich drafts, a consistent voice and a coherent explanation outweigh the number. The fair outcome is no further action, with a short note recording why the flag was cleared.
- Score high, evidence thin either way. No drafts survive and the conversation is inconclusive. Absence of drafts is not proof of guilt, so the proportionate outcome is no formal finding plus an agreement to keep drafts next time.
- Score high, with a citation or paraphrasing problem. The real issue is often referencing, not generated text. Refer the citation matter to a learning conversation and close the AI concern separately.
- Score high, several independent signals converge against authorship. A one-paste timeline, fabricated references and an explanation the student cannot sustain point the same way. Only here does a formal referral become proportionate.
- Score low but the work still looks wrong. A low score never clears misconduct on its own; weigh the same process evidence, because the number is only one signal in either direction.
How to invite a wrongly-flagged student to defend their work
Tell the student plainly that a signal raised a question, that it is not a finding or an accusation, and that their own evidence is what will settle it. Give notice, allow time, and, where policy permits, let them bring a support person. A student who genuinely wrote the work can usually defend it with concrete material, so make it easy to produce.
- Share the drafts and version history from the writing app or document, including autosaved revisions and comment threads.
- Produce earlier authentic coursework so voice, level and habits can be compared side by side.
- Walk through the sources, the argument and the specific choices made, including dead ends and passages that were cut.
- Explain any research notes, outlines or reading lists that show the work forming over time.
- Offer to redo a comparable short task under supervision if the policy allows it.
Does simple, clear, or heavily edited writing raise a detection score?
Yes. Clear, simple, conventional prose and heavily edited text all tend to read as more statistically predictable, which is what pushes a detection score up. Polished, careful human writing can therefore score higher than rough, uneven drafting.
A student who simplifies sentences for clarity, follows a textbook structure, or runs several rounds of tidying edits is smoothing out the very irregularities a detector reads as human. That is one more reason a high number is not evidence of who wrote the text, and a reason to weight the editing trail and earlier work above the score. Our explainer on why detectors flag writing a person actually wrote works through the mechanism.
Can a high score alone justify a misconduct referral?
No. A high score tells you the text reads as statistically predictable; it does not tell you a machine produced it, and it carries no information about who typed it. On its own, a high score justifies a closer look and a fair conversation, nothing more; a referral needs corroborated process evidence standing behind it.
A detection score must never be the sole basis for an academic integrity accusation. A number that estimates textual predictability is not evidence of authorship. Without corroborating process evidence and a fair conversation, a score warrants investigation and nothing further.
Documenting the case: a sample record
A neutral written record shields the student from an arbitrary decision and shields you and the institution from a decision no one can reconstruct months later. Keep it factual. Here is what a short, defensible record looks like.
Signal: detector returned 88% on Essay 3, noted 14 May. Evidence: version history showed steady drafting over six days; two cited sources could not be located in the library catalogue. Conversation: 16 May, student and module lead present; the student explained the argument and reading clearly but could not identify the two missing sources. Human review: second marker agreed the citation problem was substantive and the AI concern was not. Outcome: no misconduct finding on AI use; the citation issue referred to a learning conversation. No step relied on the score alone.
Every line records an observation, not an adjective. "Two references could not be located and the student could not identify them" is defensible; "the essay felt like AI" is not.
When the student writes English as an additional language
Yes, writing English as an additional language raises the risk of a flag, so it should lower your confidence in the score rather than raise your suspicion of the writer. Fluency is a reason for caution about the number, never grounds for an accusation.
Clear, conventional prose is exactly what pushes a predictability score up, and a student who learned English from textbooks and academic models often writes in that measured, well-structured way. A detector reads that fluency as machine-like. If a student's voice, sentence habits and level stay consistent with their earlier authentic coursework, the score is telling you about their style, not their honesty.
Verification and next steps
Before you record any outcome, run this final check. If you cannot answer yes to every line, the process is not finished.
- Did you gather process evidence, drafts and version history, and not just a score?
- Did the student get a genuine, non-accusatory chance to explain their work?
- Did a second person share the human review?
- Is the conclusion proportionate to the corroborated evidence, using the outcomes ladder above?
- Could you explain the decision to an impartial colleague using only what you wrote down?
When the concern is really weak referencing, a teaching response often fits better than a disciplinary one; our overview of whether Turnitin can detect ChatGPT shows how much varies between systems, and a practice tool such as the AI Essay Grader lets a student self-check structure before resubmitting. A fair process hands you not certainty but a defensible, humane decision that respects both academic integrity and the person in front of you.
Frequently asked questions
Can an AI detection score alone prove a student used AI?
No. A detection score estimates how statistically predictable finished text is, not who wrote it. It can prompt a closer look, but never serve as the sole basis for an academic integrity finding, which needs corroborating drafts, version history, a conversation and a second reviewer.
I am worried I will wrongly accuse an innocent student. How do I avoid that?
Treat the score as a question, not an answer, and require corroborating process evidence plus a non-accusatory conversation before any conclusion. Because the procedure forces you to look for signs the work is authentic and to involve a second reviewer, a hunch cannot become a finding.
What evidence should I weigh alongside a detection score?
Weight process evidence most heavily: drafts, version history, document metadata, and comparison with the student's earlier authentic work. Their explanation of sources and the state of their citations add corroboration. A score alone is weak; several independent signals pointing the same way support a fair decision.
Why are non-native English writers flagged more often?
Detectors score predictability, and clear, conventional, textbook-structured prose looks predictable to a language model. Writing English as an additional language often follows those patterns, so fluent, careful non-native writing is disproportionately flagged. That lowers confidence in a high score rather than raising suspicion of the writer.
How should I document a suspected AI-use case?
Record the initial signal and its date, the evidence gathered and what each item showed, the substance of the conversation in neutral language, who conducted the human review, and the conclusion with the corroborated evidence it rests on. Note observations rather than impressions, so the decision can be explained fairly later.