The Problem With Proving You Wrote It
On AI detectors, false certainty and the burden of proving that your own work is yours.
There’s a strange intimacy to writing on Substack that I love.
We spend months reading someone’s thoughts, learning what they notice and recognising the questions they return to, without always knowing very much about what they do when they close the tab.
When I close one, I generally open another.
Sometimes it’s a blank page where I draft a new essay or work through a new idea. Other times, the words belong to someone else - a member of our team or a writer I’ve commissioned.
That has been the rhythm of my working life for more than fifteen years.
For those who don’t know, I’ve always worked with words.
Copywriter. Strategist. Editor. Managing editor. I’ve worked across almost every part of content. While the work itself has expanded over the years, it’s important to know that all of that experience came before generative AI.
I share that to say, I’ve commissioned hundreds, if not thousands of articles from writers all over the world, and I still care far more about whether the writing is accurate, emotive and worth publishing than whether a monetised detector thinks AI was anywhere near it.
Perhaps that’s why it’s felt so strange to watch writers here begin proving that their work is their own. Not because they wanted to explain how they write, but because suspicion had made the explanation necessary.
What was it actually detecting?
I have used AI detectors in the past. As an editor, there was a point when it became a requirement. One we quickly shot down when we realised what it was really checking for. Patterns.
We tested writing we knew was original, including our own, and watched it come back as AI. The tools were responding to predictability, repetition, sentence structure and consistency. All things that can appear in generated text, but also in the work of people who have spent years learning to write clearly.
The result couldn’t tell us who had written the piece or how it had been produced. It could only tell us how closely the language matched the patterns the detector had learned to associate with AI. The same patterns we’d associated with good writers over the years.
It was a tool, but it wasn’t helpful. All AI detectors can do is measure how the writing looks, while our job was to understand what the writer had done with the assignment and whether they had fulfilled the brief.
Mine are usually obsessively detailed, so the subject, structure and questions are clear before they begin. What I can’t brief is the thinking. Whether they understand why the piece matters, read beyond the most obvious sources, notice something I hadn’t thought to ask about or make decisions that strengthen the work rather than simply filling the space beneath each heading.
Fluency helps, of course, but it never told me everything. Some of the cleanest submissions followed every instruction and still said very little. Some of the more awkward pieces needed work, but contained an observation that made the entire piece better.
That’s one of the things I love about editing. Even with a detailed brief in front of them, you can never predict exactly where the writer will take it.
After reading human words this closely for so long, you’d think I’d feel confident identifying AI writing. I don’t. If anything, those years have made me less certain that there was ever one recognisable way for a human being to sound. But I do recognise good writing.
Long before ChatGPT, writing was already shaped by language, education, confidence, profession, circumstance and whatever was happening in someone’s life when they sat down to begin.
I understand why we want a way to distinguish thoughtful human work from language somebody generated and submitted without understanding. Teachers want to protect the students who completed the work honestly. Editors and readers want to know that somebody was genuinely present in what they’re reading. I want to know when the thinking has been handed over too.
I’m simply not convinced that the finished paragraph can tell us.
There has never been one way to sound human
Some people are naturally repetitive, while others are concise.
There are writers who favour longer sentences, and writers who express themselves in shorter, more contained ones. One person relies heavily on transitions. Another moves between ideas so abruptly that an editor has to build the bridges for them.
A draft might be dictated, developed slowly across a week or typed at the speed of whatever the writer is trying to understand. It might have been translated, heavily edited or shaped through a conversation with somebody else. The final version rarely tells us how it began.
Assistance is woven through most professional writing now. Grammarly sits behind a great deal of it. One person asks Claude to make a paragraph clearer, another discusses possible structures with ChatGPT before writing every word themselves, while someone else translates an idea from another language and works backwards until it sounds right in English.
I use AI too. Sometimes it helps me see an argument from another direction or notice where I haven’t explained myself clearly enough. Sometimes it offers an edit that is technically smoother and completely wrong for the way I write. I can read the suggestion, understand why the model made it and still know that I would never have chosen that sentence myself.
That judgement is easier to make after years of writing. It must be much harder for someone who is still learning to trust their own voice, especially when the suggestion comes with the confidence of a system that has been presented as more objective than they are.
Plenty of people avoid using AI entirely and still produce work that detectors classify as machine-generated (with 100% ‘certainty’). Others use it extensively and pass without attracting attention. That doesn’t suggest a clear line that detectors occasionally struggle to find. It makes me wonder whether the clean boundary they require was ever even there.
Human writing has always been shaped by education, geography, class, disability, profession, personality, convention, software, collaboration and the expectations of whoever will eventually read it. A student writing a dissertation sounds different from the same person texting a friend. An SEO writer who has spent a decade answering search queries clearly and efficiently may have developed patterns that now resemble language-model output, partly because those models learned from writing produced for the same internet.
A detector still needs a baseline. It has to form some statistical idea of what human writing usually looks like and compare the submitted text with it.
The difficulty begins with the word ‘usually’.
Somebody has to decide which writing represents the centre. Some forms of vocabulary, rhythm and structure will sit comfortably inside it. Others will be further away. Once difference is translated into probability, the people whose language falls outside that centre become more likely to look suspicious.
The detector may have found a difference.
That difference alone can’t tell us who wrote the text.
The page only shows what survived
This is what makes Pangram, and the wider category of AI detection, so interesting to me. I don’t think the central problem is that the engineers behind these systems are uniquely careless or that they simply need a larger dataset. They’re trying to answer a question that the finished text may not contain enough information to answer.
Authorship isn’t visible inside a paragraph. Neither are intent, drafts or the thinking that produced them.
We can’t see the conversations that shaped an idea, the books someone read years ago, the notes they made at two in the morning or the sentences they deleted the following day. The final version doesn’t show us the argument they abandoned, the editor who challenged them or the friend who pointed out that something didn’t make sense. It can’t tell us whether they accepted an AI suggestion, rejected one or spent half an hour arguing with a chatbot before realising what they thought.
The detector receives the finished paragraph and is asked to reconstruct everything that happened before it.
Even the word ‘detector’ creates a reassuring impression. It suggests that something concealed inside the writing is waiting to be found. A plagiarism checker can compare a sentence with existing material and identify a match. A spellchecker can recognise a word that falls outside its dictionary. In those cases, the system is searching for something present in the text.
AI detectors are doing something more uncertain. They examine statistical characteristics and estimate which category the paragraph resembles. They don’t recover the history of the work or observe the decisions that produced it. They create a probability from the surface that remains.
That may still be useful as one small signal among many. The problem begins when the probability is asked to carry more weight than it can bear, particularly when the person receiving it is already under pressure to reach a conclusion.
A lecturer may have hundreds of essays to mark and a genuine responsibility to protect the students who completed the work honestly. An editor may be screening more submissions than anybody could read properly in the time available. A reader may have subscribed to someone because they valued the feeling of hearing directly from them and now worries that the relationship is being quietly automated.
I understand why a number feels helpful in those moments. Uncertainty is exhausting, especially when getting the judgement wrong seems unfair in every direction.
The score appears to offer somewhere firm to stand.
The test is cleaner than the writing
Companies selling AI detection often advertise accuracy rates in the high nineties. Those figures may be accurate within the conditions of the test, but the conditions matter. Cleanly generated AI text is compared with cleanly human-written text, often using subjects, models and forms of writing already represented in the evaluation data.
Very little writing remains that clean once people begin working with it.
People edit. They paraphrase. They combine their own sentences with generated ones. They accept one suggestion and reject another. They move sections around, translate passages, shorten them and pass them between tools. Even when someone generates an entire first version, the piece that eventually reaches an editor or lecturer may be several rounds of human thought and revision removed from it.
Research reflects that gap. On paraphrased, edited or mixed text, detector performance can fall by roughly 25 to 50 points. A separate evaluation of seven popular detectors found 39.5% baseline accuracy on unaltered AI text, dropping to 17.4% after relatively simple adversarial editing.
Those conditions are often described as difficult. They also resemble the way a great deal of writing is now made.
Accuracy is only one part of the question. Reliability asks whether the same tool produces consistent results when nothing meaningful has changed. A study of 14 AI detection tools concluded that they were neither accurate nor reliable. Every system scored below 80% accuracy and only five exceeded 70%, with human work classified as AI and AI-generated work classified as human.
A result that changes according to the detector, model version or testing conditions remains a prediction. That uncertainty needs to stay visible when somebody’s grade, work or reputation may depend on what happens next.
The RAID benchmark brought more of the problem into view. Researchers assembled more than six million generated texts across different models, subjects and forms of manipulation, and found that detectors were easily disrupted by unfamiliar models, generation settings and relatively straightforward changes to the writing.
The false-positive question matters most to me because every false positive belongs to a person.
It’s easy to discuss percentages at a distance. It feels different when the percentage becomes a student sitting outside a misconduct meeting, trying to remember whether they kept enough drafts. Or a writer opening an email from a client who suddenly doubts work they spent days producing. Or somebody writing in a language they learned later in life, being asked to explain why their vocabulary and sentence structures look unusual to a machine.
A detector can appear more effective if it is allowed to accuse more human writers along the way. In any setting where the output may be used against someone, keeping that number extremely low should be the minimum condition for using the result at all.
There’s also a limit that better engineering can’t remove. Researchers have shown that even an ideal detector is constrained by the statistical distance between human and machine-generated language. As the two become more similar, reliable classification becomes mathematically harder.
That distance is shrinking from both directions. Models continue improving at imitating the range and irregularity of human writing, while human language absorbs patterns from systems we have now been reading and working with for several years.
There may never have been a stable fingerprint of human writing. We’re making the idea of one even less plausible.
The uncertainty is already in the small print
What interests me is how little of this uncertainty is hidden.
Turnitin’s own guidance says its model may misidentify human, AI-generated and AI-paraphrased writing, and that the result shouldn’t be used as the sole basis for adverse action against a student. Originality.ai describes its results as probabilistic estimates affected by language, genre, translation, editing, paraphrasing, text length and mixed authorship. Its terms prohibit using them as the sole basis for decisions that could significantly affect a student, employee, applicant or writer. Pangram’s technical report includes an ethics discussion cautioning against treating its classifier as the sole arbiter of academic integrity.
The people building these tools have written the limitations down. False positives are possible. Human review remains necessary. The output shouldn’t be used alone.
And then the software enters institutions because those institutions are desperate for help making the very judgement the disclaimer says the tool can’t make for them.
I don’t think that contradiction is always driven by carelessness. Universities and publishers are trying to respond to a genuine change in how work is produced, often without extra time, staff or useful guidance. Lecturers are expected to preserve academic standards while distinguishing between dozens of different kinds of assistance that may leave no visible trace. Editors are trying to protect trust with readers while working inside business models that increasingly reward speed.
The score arrives at exactly the moment everyone feels least able to remain uncertain.
In the UK, the sector’s advice has been cautious for several years. Jisc, the organisation advising British universities on technology, says no AI detector can conclusively prove that text was written by AI and that it can’t recommend any current product as reliable enough. It also advises lecturers against uploading student work to unapproved third-party systems, partly because the students haven’t consented to it.
The University of Greenwich publicly opted out of Turnitin’s AI detection feature, explaining that it couldn’t assess the accuracy, the likelihood of false positives or the potential effect on student outcomes.
The consequences those warnings are trying to prevent are already documented.
In July 2025, the Office of the Independent Adjudicator, the ombudsman for higher education in England and Wales, published several case summaries involving students accused of AI misconduct.
In one case, a student explained during a misconduct viva that they had used Google to find synonyms because English wasn’t their first language. The panel recorded this as an admission of using AI to paraphrase. The OIA partly upheld the complaint and found, among other problems, that the university hadn’t considered whether detection software might perform less reliably for people writing in a second language.
I keep thinking about what that conversation must have felt like from the student’s side of the table. To be explaining how you searched for a word in a language you are still learning, while the people questioning you interpret the explanation through a suspicion that existed before you entered the room. Even an ordinary attempt to communicate becomes something that can be used to strengthen the accusation.
In another case, an international student wasn’t shown the evidence against them before their viva, and their explanation that they had used Grammarly was dismissed without proper consideration. Their complaint was upheld in full.
The third case involved an autistic student who was accused, penalised and later cleared when the decision was reconsidered. An earlier submission by the same student had also been flagged by the software and found to be entirely human-written.
Twice, the tool treated the way that person wrote as evidence of possible misconduct. Twice, no AI had been used.
There’s an obvious institutional failure in those cases, but I also think there is something more ordinary and human happening underneath it. Once a score creates suspicion, it is difficult for the people involved to return to the neutrality they had before they saw it. Every explanation begins to sound like a defence. Every unusual choice can be interpreted as further evidence.
A student may have drafts or document histories. A writer may be able to explain how they work. Yet their account has to compete with a number that appears calmer, cleaner and less emotionally invested than anybody in the room.
The system produces a probability. People then begin building a story around the person.
A model of normal always leaves someone outside it
The harms created by that process won’t be distributed evenly.
In the most widely cited study on this question, Stanford researchers ran essays written by non-native English speakers through seven commercial detectors. On average, 61% were flagged as AI-generated. Essays written by US schoolchildren aged 13 and 14 were identified as human with near-perfect accuracy. Every one of the TOEFL essays had been written by a person.
What the researchers found underneath the numbers matters even more. The essays being flagged tended to have less linguistic variability. When a language model rewrote the same work using richer vocabulary and more varied sentence structures, the bias disappeared.
The detectors were responding to a narrower vocabulary and more consistent structure, qualities that already exist across a great deal of human writing.
People writing in a second language may rely on vocabulary and structures they know they can use accurately. Academic writers are often taught to follow strict conventions. Technical and SEO writers develop clear, repeatable structures because their work requires precision rather than stylistic surprise. Autistic writers may use repetition, explicitness or forms of logical organisation that sit outside whatever range the system has learned to associate with human expression.
None of those qualities make the writing less thoughtful, less original or less human. They may simply make it less similar to the people whose work sits closest to the detector’s idea of normal.
There is also an instruction hidden inside the Stanford findings. If richer vocabulary and more varied sentences make the suspicion disappear, the system is doing more than evaluating writing. It’s teaching people how they need to write if they want to avoid being questioned.
Use a less obvious word. Vary the rhythm. Break the structure that looks too orderly. Add enough irregularity to become recognisable as a person.
I don’t know how many people are already changing their writing this way, consciously or otherwise. I imagine some are doing it simply because they are frightened. Once your own writing has been treated as evidence against you, running the next piece through a detector before submitting it may feel like the sensible thing to do.
Then a strange reversal begins. A person writes something entirely themselves, receives a suspicious score and alters it until the machine agrees that it looks human. The accepted version may be further from their natural voice than the one that failed.
That possibility deserves its own conversation. What happens to a body of writing when everyone inside it knows they may be scored? How many people will begin composing for the classifier rather than the reader, and how much of the real variety of human expression will slowly be edited away in pursuit of the green light?
Every algorithm has to build a model of normal. The people furthest from that model are then asked to explain why they are different.
What the writing does matters
When I commission writers, I’m evaluating whether the piece does what it was meant to do.
That means looking at whether they followed the brief, understood the tone and thought about the person who would eventually read it. Whether the piece feels useful, relatable and right for the publication. Whether it answers the question it set out to answer without losing the reader somewhere along the way.
Accuracy is important, of course, but it’s never been the only thing I’m looking for. A piece can be technically correct and still feel distant. It can follow every heading in a brief and still miss the reason somebody would want to read it.
I’m also paying attention to the decisions inside the work. The examples a writer chose. The point they decided to spend more time on. Whether they noticed something beyond the obvious answer or understood what the reader might be worried about, curious about or hoping to find.
I’ve received impeccably written submissions that followed every instruction and still left me cold. There was nothing obviously wrong with them. They simply didn’t make me feel, notice or understand anything differently.
I’ve also commissioned writers whose submissions needed work, but contained an observation I hadn’t come across before. The sentences could be improved without losing the reason they were worth reading. A grammatical mistake could sit beside an insight I hadn’t considered, while a perfectly structured article could leave me with no sense that the writer had thought about the person on the other side of it.
That is why editors read.
We’re evaluating accuracy, judgement, tone, relatability and how the work feels to encounter as a reader. I don’t care whether somebody used Claude to fix a comma. I care whether they understood the brief, whether their choices make sense and whether the finished piece does what it was meant to do.
Those qualities don’t leave a reliable statistical fingerprint. They become visible through the experience of reading the work.
Human judgement isn’t infallible. It carries bias, inconsistency and its own ideas about what good writing should sound like. I have mine.
There are submissions that feel wrong to me before I can fully explain why, and I’m not confident the problem is always the writing. It may be the register, a formality I interpret as distance or a structure so orderly that I assume nobody has struggled inside it.
At times, I’ve probably read texture and called it judgement.
So the instinct being automated isn’t foreign to me. I know what it is to read a piece and feel uncertain about the person behind it. I also know that feeling doesn’t make my interpretation true.
The difference is that I don’t place my uncertainty inside a number and ask it to speak for me. If I turn something down, I have to decide what I believe is missing. If the decision matters to somebody, I have to be willing to explain it and accept that I may have read them wrongly.
A human judgement can be questioned. It can hear context, reconsider the evidence and apologise when it gets something wrong.
A score offers more distance.
Why certainty feels so comforting
Reading properly is slow.
It requires attention, knowledge of the person or subject and a willingness to remain uncertain while gathering more information. It may require a conversation with a student or writer. It may end without the satisfying certainty everyone hoped to find.
Someone also has to be accountable for the judgement that follows.
A score seems to reduce that burden. The lecturer didn’t accuse the student, the detector flagged the work. The publisher didn’t distrust the writer, the software returned a probability. The recruiter didn’t overlook the applicant, the system ranked them lower.
I understand the appeal, and I don’t think the people reaching for these tools are lazy or indifferent.
The lecturer marking 200 scripts to a deadline didn’t design that workload. They may be worried that students who completed the work honestly are being disadvantaged by those who didn’t. They may also be frightened of accusing somebody unfairly and have been given very little practical help navigating the space between those two risks.
An editor screening hundreds of applications is rarely given enough time to read every one closely. A publisher watching readers lose trust in online writing may feel that doing nothing is no longer an option. A reader who has supported a writer financially may want reassurance that the person they came to hear is still present in the work.
The desire for certainty grows from understandable places.
The problem is that the score can’t provide the kind of certainty we are placing inside it.
It can give uncertainty a shape. It can turn a complicated concern into a percentage and make action feel possible. What it can’t do is remove the need for judgement. Someone still has to decide what the result means, how much weight it deserves and what should happen to the person on the other side of it.
Human review is often described as a safeguard around the score. I think it should be the beginning of the process, because the human context is precisely what the detector can’t see.
The software can notice a pattern. It doesn’t know whether a student searched for synonyms because they were trying to express themselves more accurately in a second language. It can’t understand that repetition may be part of an autistic person’s natural communication. It can’t tell whether a writer accepted one suggestion, rejected twenty others and remained responsible for every decision that survived.
It also can’t carry the responsibility for an accusation.
That remains with us, however appealing the distance of the number may be.
Generative AI creates real problems. It can be used to avoid work, conceal a lack of understanding and submit language that someone neither wrote nor meaningfully shaped. Schools, publishers and employers need ways to respond to that honestly, and the people doing that work deserve more support than a disclaimer wrapped around a confidence score.
I am simply no longer convinced that the most useful question is whether a machine has been somewhere near the words.
Perhaps we should be asking whether the person understood what they submitted, whether they can explain the decisions inside it and whether there is evidence that they did the thinking the work was meant to require.
Those questions take longer. They leave more room for uncertainty, and they ask something of the person making the judgement too.
They also give the person being judged a chance to be heard.
The score can’t tell us who was thinking.
It can only make us feel, for a moment, as though we no longer need to ask.


I don’t feel that I, or anyone writing here, has anything to “PROVE”. And the notion that there is now an unhuman machine that will judge what we write disgusts and enrages me. The only way I will be able to continue to engage with Substack is by IGNORING this whole Pangram thing and all the articles about it. Fortunately I’m pretty good at blocking out stuff that annoys me!
I would love see how these tools evaluate texts from Jane Austin, Ernest Hemingway, Annie Dillard, P G Wodehouse, etc.