AI detection cannot carry the weight faculty need it to carry
Detectors report a probability about how text was produced. Faculty have to make decisions that need evidence they can inspect and explain, and a score supplies none of the process that makes a decision defensible.
The Provenance Team
Faculty need reliable evidence. That is the real issue.
When generative AI became widely available, AI detection looked like the obvious institutional response. If students could use tools to produce fluent text, then perhaps other tools could identify that text and restore the old balance. Faculty could check a report, confirm whether AI was likely involved, and move forward with confidence.
That hope was understandable, but it was too heavy a burden for detection to carry.
AI detection tools estimate likelihood. They do not observe the student’s process, they do not know what happened before submission, and they do not explain the student’s understanding of the work. A detector may flag text as likely AI-generated, but the report usually cannot tell faculty which choices the student made, what sources the student used, what they revised, what they understood, or whether the submitted artifact reflects the learning outcome.
That distinction is practical, not philosophical. Faculty are asked to make decisions about grades, course credit, degree progress, and sometimes academic integrity referrals. A probability score may raise a question, but a serious academic decision requires evidence that can be reviewed, explained, and connected to the course criteria.
Detection rarely gives faculty that kind of evidence.
The tools do not claim to be proof
Whether detection is accurate is genuinely contested, and this piece does not try to settle it. What matters for an institution is less disputed: the tools themselves do not claim to be proof. OpenAI released an AI text classifier in early 2023 and withdrew it months later, citing low accuracy (OpenAI, 2023). Turnitin, whose detector remains widely used, cautions that false positives are possible and treats low-percentage results with particular care, since false positives are more common in that range (Turnitin, 2024). Turnitin also states that its score is meant to inform an educator’s judgment, not to decide misconduct on its own.
That vendor caution is the point. When the makers of these tools say a score is a signal rather than a verdict, an institution should not build a verdict on it. Detection depends on statistical patterns in language, and reasonable people continue to disagree about how well it performs across genres, editing, translation, and writing level. An integrity regime that has to win that argument in order to function is already standing on unstable ground.
Detection invites fairness disputes
Fairness adds another problem, and it is one institutions increasingly meet in appeals and courtrooms rather than in the abstract.
Detectors work by spotting statistical patterns in language, so the writing most likely to be misread as machine-generated is the writing that already looks predictable or formulaic — and that describes a great deal of legitimate human work. Researchers have raised this concern specifically for non-native English writers, whose prose can register as lower-variability to a detector (Liang et al., 2023); how strong that particular finding is has since been debated, and some newer detectors report better results, but the underlying risk does not disappear. Many institutions serve multilingual students, adult students, neurodivergent students, students using writing support, and students whose prose may be direct, formulaic, highly edited, or uneven for reasons that have nothing to do with prohibited AI use.
A tool that mistakes some kinds of human writing for machine writing can do real harm, even when no penalty is imposed immediately. The student is placed under suspicion. The faculty member inherits a difficult conversation. The institution now has a record that may be hard to interpret, and if the process is not handled carefully, the student may experience the assessment system as arbitrary rather than fair. Whether or not any single flag is correct, the perception of bias is now a routine feature of AI-era integrity cases — and perception is enough to produce disputes, appeals, and reputational cost.
The issue, then, is not whether detection tools are useless. They are not. In some cases, they may help identify work that needs a closer look. They may support a conversation, prompt a review, or contribute to a larger evidence set. Used carefully and with institutional safeguards, they can play a limited role.
The real danger: detection as strategy
The danger comes when detection becomes the assessment strategy.
A detector report can shift attention away from learning and toward production history. Instead of asking whether the student can explain the argument, interpret the data, walk through the code, or justify the design, the faculty member is left asking whether the text pattern crosses some invisible threshold. That question may be relevant in a policy process, but it does not tell us enough about the student’s learning.
It also puts faculty in a bad position. Many instructors can recognize when something feels off, yet they also know that intuition is not enough. They may compare the submission with earlier writing, run the work through a detector, inspect citations, check document history, and request a meeting with the student. Even then, the evidence may remain thin. The process becomes time-consuming, emotionally charged, and difficult to defend if challenged.
Courts are now making that lesson concrete. In January 2026, a New York court annulled an Adelphi University misconduct finding that rested on a Turnitin result reported at 100% AI-generated (Matter of Newby v. Adelphi University, 2026). The court did not rule that the detector was wrong; it never reached that question. It ruled that the university had acted arbitrarily — treating the score as proof while denying the student a meaningful opportunity to be heard, disregarding the contrary evidence he submitted, and letting the same official decide both the charge and the appeal. The finding was annulled and ordered expunged. The lesson is not that detection failed. It is that a score cannot carry a misconduct decision without a fair, reviewable process around it, and a detector supplies no such process.
The more useful question is different: what would the student need to explain for me to judge whether the learning outcome was met?
For a research paper, the student might need to explain why one source was more credible than another, how the thesis changed during revision, or what evidence would weaken the argument. For code, the student might need to trace a function, explain an error-handling choice, or describe why one data structure was selected over another. For a case analysis, the student might need to connect a recommendation to a specific case fact, identify an alternative interpretation, or explain the risks of the proposed solution.
These questions produce evidence that detection cannot provide. They show whether the student can reason from the artifact, not merely whether the artifact contains suspicious textual patterns.
Shift to the review surface
That is why the submitted artifact should become the review surface.
The artifact is still central. Faculty should not abandon authentic assignments because AI makes authorship harder to infer from the final product alone. The better move is to make the artifact more inspectable. Students should be asked to revisit their submitted work, point to specific sections, explain choices, and demonstrate understanding in relation to the criteria faculty already approved.
This kind of review changes the faculty task. Instead of relying on a detector to infer production history, the instructor can inspect the student’s explanation alongside the artifact. The record can include the question asked, the part of the artifact referenced, the student’s answer, the transcript, the timestamp, the rubric connection, and the faculty’s final judgment.
That is a stronger evidentiary base.
Where Provenance fits
Provenance Learning starts from the opposite premise to a detector. Instead of estimating how a text was produced, it helps faculty see what a student understands: it generates and approves questions grounded in the assignment, rubric, outcomes, and the submission, then lets students respond by voice or text. The session captures audio, timecoded transcription, and artifact references, and produces evidence-linked recommendations faculty can override, with faculty holding the criteria and the final results. A detector hands back a probability about the past; this hands faculty an inspectable account of what the student can explain.
This approach does not require Provenance to determine whether misconduct occurred. It should not make that determination. Academic integrity processes belong to institutions and faculty governance structures, and any system that touches student evaluation should respect those boundaries.
What Provenance can do is produce better evidence.
If a student understands the artifact, that understanding should become visible. If the student cannot explain central claims, design choices, calculations, sources, or code, that gap should also become visible. In either case, faculty should be able to review the record and make a human academic judgment.
AI detection asks a narrow question about how text may have been produced. Faculty need a broader answer about what the student understands.
The difference is not cosmetic. It changes the evidence, the student experience, the faculty workload, and the defensibility of the decision.
Detection may have a place, but it cannot be the center. The center has to be demonstrated understanding.
Key takeaways
- AI detectors report probabilities, not proof — and the vendors themselves say a score should inform an educator’s judgment, not replace it.
- Detection invites fairness disputes and a perception of bias, especially for multilingual, disabled, and neurodivergent students — concerns detection cannot resolve on its own.
- The real risk is treating a detection score as a verdict; in 2026 a New York court annulled a misconduct finding built on a 100% score, faulting the school’s process rather than the tool.
- Artifact-based review gives faculty an explanation they can inspect instead of a score they cannot.
References
Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), Article 100779. https://doi.org/10.1016/j.patter.2023.100779
Matter of Newby v. Adelphi University, 2026 NY Slip Op 26021 (Sup Ct, Nassau County Jan. 28, 2026). https://law.justia.com/cases/new-york/other-courts/2026/2026-ny-slip-op-26021.html
OpenAI. (2023, January 31). New AI classifier for indicating AI-written text. https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
Turnitin. (2024, July 18). AI writing detection model. https://guides.turnitin.com/hc/en-us/articles/28294949544717-AI-writing-detection-model
Editorial note
This article was drafted with AI assistance and reviewed, revised, and approved by Provenance Learning’s human authors. The claims, sources, examples, and final wording were reviewed before publication. We believe responsible AI use should be transparent, reviewable, and subject to human judgment. See our Editorial Transparency statement.
Provenance is an artifact-based oral assessment platform for higher education. If you’d like to see a session run on one of your own assignments, request a demo.