When the paper trail can be counterfeited
Version histories and process logs looked like a sturdier alternative to AI detection — until tools appeared that manufacture them. A process trace holds up better when it is tied to explanation than when it is treated as proof on its own.
The Provenance Team
Faculty have spent the past few years looking for better evidence of student work. If a finished paper, project, case analysis, or code submission no longer tells us enough on its own, then drafts, version histories, reflections, conferences, and process logs seem like reasonable places to look.
They are useful, but a recent Chronicle essay by Susan E. Ray shows why we should be careful about what we think they prove.
Ray describes testing a set of tools marketed directly to students: an AI “humanizer” that rewrites generated text to avoid detection, followed by an auto-typer that enters the finished essay into Google Docs slowly enough to resemble ordinary drafting. Over several hours, the software inserted text, paused, made errors, corrected them, and created a version history that looked remarkably like a student writing over time (Ray, 2026).
The result is unsettling because version history has seemed like one of the more sensible alternatives to AI detection. Faculty can inspect how a document developed instead of relying on a probability score about its final prose. Instructors have encouraged students to work in shared documents, preserve drafts, maintain AI-transparency journals, and document revision partly because those records make the learning process more visible.
Ray’s experiment exposes a weakness in that approach. A version history records activity in a document; it does not establish authorship or understanding. For years, the distinction was easy to overlook because there was little reason to simulate hours of typing. Now there are products designed specifically to do so, and not only paid ones. At least one free, open-source browser extension does the same job, written to work in Google Docs and in common course platforms. Manufacturing a plausible writing history no longer costs a student anything.
The pattern is becoming familiar. Detection tools are followed by humanizers. Concerns about large blocks of pasted text lead to version-history checks, followed by auto-typers that manufacture gradual entry. Faculty identify a signal that appears useful, and a product emerges that can imitate or obscure it.
This does not make process evidence worthless. Drafts, conferences, revision histories, source logs, and transparency journals can still tell faculty a great deal, especially when they are part of an ongoing course relationship. Ray herself reaches a similar conclusion: the strongest evidence in her own teaching comes from the accumulated record of student work that she has witnessed over time—drafting, feedback, revision, reflection, and growth (Ray, 2026).
The problem comes when one trace is expected to settle the question.
Research on AI detection has already shown the danger. Weber-Wulff et al. (2023) tested 14 detection tools and found them “neither accurate nor reliable,” with a main bias toward classifying AI-generated output as human-written, and with performance falling further after obfuscation techniques such as manual editing or machine paraphrasing. Their conclusion was direct: the tools are unsuitable for use as evidence of academic misconduct, and their reports cannot be the only basis for reporting a student (Weber-Wulff et al., 2023). Meanwhile, institutions are spending heavily on this category of technology; CalMatters and The Markup documented more than $15 million in Turnitin purchases across 57 California institutions (García Mathewson, 2025).
Faculty can end up caught between two markets. One sells institutions increasingly sophisticated ways to inspect student work, while another sells students increasingly sophisticated ways to manufacture the signals those systems are looking for. The cycle produces more telemetry, more uncertainty, and more faculty labor.
A stronger assessment practice asks a question that these tools cannot answer very well: what does the student understand about the work they submitted?
A version history can still help generate useful questions. Why did this section change? What led you to replace this source? How did your argument develop between the outline and the final draft? If AI use was permitted, what did the tool contribute, what did you reject, and what did you verify yourself?
The submitted artifact can support the same kind of review. A student who submits code can explain a function, trace the data flow, or identify a likely failure case. A student who writes a research paper can explain why one source was more persuasive than another or what evidence would weaken the central claim. A student who produces a case analysis can explain why particular facts drove the recommendation and where the proposed solution is vulnerable.
These questions move the evidence closer to learning because the student has to account for specific features of the artifact.
This is the premise behind artifact-based oral assessment. The artifact becomes the review surface: students revisit the work they submitted, faculty-approved questions point to specific parts of it, and the resulting explanations become part of the evidentiary record. A transcript, timestamp, artifact reference, rubric connection, and faculty judgment can then sit alongside the original submission.
Provenance Learning is being built around that model. Faculty retain control over the questions, criteria, and final decision, while the system helps conduct and document the review. It does not determine misconduct, and it does not attempt to infer authorship from textual or behavioral signals. Its purpose is to help faculty gather clearer evidence of demonstrated understanding.
Ray’s experiment is useful because it shows how quickly a seemingly reliable signal can become unreliable once people have an incentive to counterfeit it. Version histories still have value, as do drafts, process journals, and other records of student work, but each is stronger when connected to explanation rather than treated as proof on its own.
The more durable question is whether the student can explain the artifact submitted in their name.
That is evidence faculty can actually use.
References
García Mathewson, T. (2025, June 26). Costly and unreliable: AI and plagiarism detectors wreak havoc in higher ed. CalMatters. https://calmatters.org/education/higher-education/2025/06/ai-detector/
Ray, S. E. (2026, August 12). The next wave of AI-cheating technology. The Chronicle of Higher Education. https://www.chronicle.com/article/the-next-wave-of-ai-cheating-technology
Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, Article 26. https://doi.org/10.1007/s40979-023-00146-z
Editorial note
This article was drafted with AI assistance and reviewed, revised, and approved by Provenance Learning’s human authors. The claims, sources, examples, and final wording were reviewed before publication. We believe responsible AI use should be transparent, reviewable, and subject to human judgment. See our Editorial Transparency statement.
Provenance is an artifact-based oral assessment platform for higher education. If you’d like to see a session run on one of your own assignments, request a demo.