← Blogacademic-integrityoral-assessment

Submitted work is no longer self-authenticating

A finished artifact no longer proves who did the thinking behind it. The better move isn't to authenticate the document — it's to let the student authenticate the learning.

The Provenance Team

For a long time, faculty could read a submitted paper, project, lab report, case analysis, slide deck, or code submission and treat it as the main evidence of a student’s learning. That practice was never without its limitations, but it functioned as a workable proxy for understanding.

The evidence was never perfect. Students could receive too much help. They could misunderstand work they had technically completed. They could submit polished writing without much judgment behind it. Faculty knew this and responded by building layered assessment structures—rubrics, drafts, conferences, presentations, reflections, peer review, and revision cycles—because finished work only ever told part of the story.

Generative AI has made that limitation harder to ignore and has shifted the burden from assumption to verification.

Students are using AI — that part isn’t in doubt

The issue is not whether students are using AI. They are, and the data confirms it. HEPI’s 2026 survey of full-time UK undergraduates found that 95% reported using AI in at least one way, while 94% said they used generative AI to help with assessed work (Stephenson & Armstrong, 2026). In a large U.S. study of 95,513 students across 20 public research universities, about one-third reported regular generative AI use for assignments involving text, video, or code, and 9% reported prohibited use (Chirikov et al., 2026).

Those numbers should not push faculty toward panic. They should push us toward better evidence and more deliberate forms of verification.

The problem is isolation, not the artifact

A polished submission can still represent real student work. It can demonstrate care, discipline knowledge, revision, judgment, and skill. AI use does not erase those possibilities, but it complicates how confidently we can attribute them. Students use AI in varied ways—for explanation, brainstorming, grammar support, coding help, study questions, and feedback. Some uses are permitted, some are not, some are educationally productive, and some hollow out the task.

The problem is not the artifact itself but its isolation. The artifact alone now carries less evidentiary weight, and without additional context, its meaning becomes less stable.

A finished artifact may show what was submitted, but it may not show what the student understands. It may not reveal which choices were theirs, what they revised, what tradeoffs they considered, what sources they trusted, what code they can explain, what claims they can defend with evidence, or what they would do differently with more time.

Faculty already sense the gap

Faculty feel this immediately. The paper reads well, but the student’s discussion posts are thin. The code runs, but one function looks unlike anything else the student has written. The case analysis uses the right vocabulary, but the reasoning feels assembled rather than owned. The reflection is fluent, but oddly generic. None of that, taken alone, is sufficient to support a serious academic decision. It does, however, raise the right question: what evidence would help me judge whether this student met the learning outcome?

Detection and lockdown can’t carry the weight

Quality agencies are asking the same question. QAA warned in 2023 that some assessments contributing to the evidence base for awards may no longer be “confidently ascribed” to an individual student when generative AI tools are available, and that AI-generated outputs cannot be reliably detected in ways that provide clear, attributable evidence (QAA, 2023). Detection methods may work in some cases, but they are widely perceived as unreliable, and because machine learning systems are opaque, they tend to produce probabilities rather than explanations faculty can point to directly. The result is uncertainty rather than clarity.

The obvious reaction is to retreat to tightly controlled exams. Sometimes that will be appropriate, because some outcomes genuinely require controlled conditions. But a wholesale return to locked-down testing would sacrifice too much. Faculty have spent years building more authentic assignments—projects, portfolios, case analyses, applied writing, presentations, labs, code, design work, policy briefs, and workplace-style deliverables. These artifacts matter because they resemble the kinds of work students will be asked to do beyond the course.

We should not discard those artifacts. We should make them more inspectable and more clearly connected to demonstrable understanding.

Treat the artifact as the review surface

That shift begins with a reframing. Treat the submitted artifact as the review surface, not as the final proof.

A student who submits a case analysis should be able to walk through a specific claim, explain why the evidence supports it, and discuss what would weaken the argument. A student who submits code should be able to explain a design choice, trace the data flow, identify a failure case, and describe what they would refactor. A student who submits a slide deck should be able to explain what was included, what was omitted, and how the intended audience shaped the structure.

This kind of review changes the evidentiary landscape. The artifact remains central, but it no longer stands alone. The student’s explanation becomes part of the record, along with questions, transcripts, timestamps, artifact references, rubric connections, and faculty judgment. Together, these create a more complete and inspectable account.

That is the space Provenance Learning is built for.

That gap — between a finished artifact and the understanding behind it — is what Provenance Learning addresses. Rather than trying to authenticate the document, it gives students a structured way to authenticate the learning: they revisit their own submission and answer faculty-approved questions tied to the assignment, rubric, outcomes, and the artifact itself. Faculty set the questions and criteria and make the final call, while the session preserves a timecoded transcript, audio, and artifact references, and produces evidence-linked recommendations faculty review and can override. The submission stops standing alone as proof and becomes the starting point for a record that can actually be reviewed.

This is not AI detection. It is not automated grading. It is not a misconduct decision engine. It is a way to turn submitted work into demonstrated understanding.

That difference becomes clear in practice. Detection asks whether a tool can infer how text was produced. Artifact-based oral review asks whether the student can explain the work that was submitted. Detection often produces suspicion without much instructional value. Review produces evidence that faculty can inspect, weigh, and validate.

The goal is not to turn every assignment into a courtroom proceeding. It is to restore a foundational academic relationship. Students submit meaningful work. Students can explain that work. Faculty have a transparent, reviewable record for judging whether learning has been demonstrated.

Submitted work is still important. In many courses, it should remain central. But in the age of AI, the submitted artifact requires reinforcement. It needs context, explanation, traceability, and faculty oversight.

Evidence, transparency, and trust.

That is the task ahead.

Key takeaways

  • A finished artifact no longer proves who did the thinking behind it; polish and authorship have come apart.
  • Detection and lockdown exams don’t restore that proof — detection is contested and only asks whether text looks AI-made, while lockdown narrows what you can assess.
  • Treating the submitted artifact as a review surface, where students explain it, turns a document into evidence of understanding.
  • Faculty keep final judgment; what changes is the evidence they have to judge with.

References

Chirikov, I., Smirnov, I., & Kizilcec, R. F. (2026). Generative AI use and misuse call for assessment reform in higher education. Science, 392, 818–820. https://doi.org/10.1126/science.aec5115

Quality Assurance Agency for Higher Education. (2023). Reconsidering assessment for the ChatGPT era: QAA advice on developing sustainable assessment strategies. https://www.qaa.ac.uk/docs/qaa/members/reconsidering-assessment-for-the-chat-gpt-era.pdf

Stephenson, R., & Armstrong, C. (2026). Student Generative Artificial Intelligence Survey 2026. Higher Education Policy Institute. https://www.hepi.ac.uk/reports/student-generative-ai-survey-2026/

Editorial note

This article was drafted with AI assistance and reviewed, revised, and approved by Provenance Learning’s human authors. The claims, sources, examples, and final wording were reviewed before publication. We believe responsible AI use should be transparent, reviewable, and subject to human judgment. See our Editorial Transparency statement.


Provenance is an artifact-based oral assessment platform for higher education. If you’d like to see a session run on one of your own assignments, request a demo.