AI visibility is the practice of recording how a student worked with AI — every prompt, every response, every revision — so a teacher can assess the process rather than estimate whether the finished text was machine-written.
A detector puts you in an argument with your student. A record puts you in a conversation.
A record of how the work was done. Three things you need in order to mark it, none of which a probability about the finished text can give you.
Which sentence, when it arrived, and what came before it. Marking a piece of writing means knowing how it was made, and that is a matter of record rather than of judgement about the finished text.
Feedback only teaches if the student can look at the same thing you looked at. A transcript of their own session is something to talk about together; it belongs to them as much as to you.
Anything that arrives after submission is a grade. Seeing a student stuck, or leaning on the AI, while there is still time to say something is the part that changes what they learn.
They are not two versions of the same tool. They ask different questions, and only one of the answers has a next step attached.
| AI detection | AI visibility | |
|---|---|---|
| The question | Might this text have been written by AI? | How did this student work? |
| What you get back | A probability about the finished text | The sequence — prompts, replies, revisions, in order |
| Where it looks | At the artifact, after the fact | At the working, as it happens |
| What you can do next | Raise it with the student, or let it go | Teach, mark against a rubric, or ask a better question |
| What the student sees | A verdict about them | Their own working, which they can learn from |
| What a parent sees | A number they are asked to trust | The exchange itself — what was asked, what came back |
| Where it struggles | Hardest on students writing in a second language | Only covers work done inside the workspace |
Accuracy figures are mostly vendor-reported and conditional on how much of a document gets flagged, so the average rate is the less useful number. What matters is who the errors fall on. When researchers tested seven commercial detectors on essays by non-native English writers, every one of them misclassified that writing as AI-generated — while classifying native-speaker writing correctly.
The likely cause is mechanical rather than malicious: a writer with a smaller working vocabulary produces more predictable text, and predictability is what these tools measure. The researchers cautioned against using them in educational settings for exactly this reason.
In a classroom where most students are writing in their second language, that failure mode lands on the same children repeatedly. Visibility has no equivalent, because it reads what happened rather than estimating it from the prose.
Liang et al., Patterns (Cell Press), 2023 — US eighth-grade essays vs TOEFL essays.
Worth saying plainly, because the last tool promised more than it delivered.
A student can draft somewhere else and paste it in. What the record shows is a page that arrived with no working behind it — which is a signal, but it is not a locked door.
Cognity does not label a student dishonest and does not produce a score you are meant to act on. It shows what happened. The judgement stays with the teacher, where it belongs.
Work done in another tab is not in the record. What Cognity can tell you is what happened in the task you set, which is the part you are marking.
Answered straight, including the ones where the honest answer is not the one we would prefer.