The Ghost in the Machine
Nobody believes the dog ate your homework.
Which is the problem, because this time it did.
Everybody worries that AI will hallucinate something.
It can, and it will — regularly. But it does other things to text as well, and the worst of them in our experience is this: it removes something true, and the prose closes over the gap.
Below are nine real examples, illustrating nine different ways these tools damage text. Each comes from our own work, with a date, a high-end frontier model, and what happened to the text.
None is hypothetical. All involved advanced frontier models — no free tools, no old versions. Some are models working directly on text. Others are models writing and running the code that edits text, which is the same failure operating at scale.
Back to Human-in-the-Loop Structural Refinement
One of nine pages on AI and academic honesty. Browse all 62.
Related links

A fabricated citation gets caught. You go to verify it, and it does not exist.
A vanished paragraph is not caught. Neither is a real source deleted as fabricated. Nothing reports either. Silently, draft by draft, the text is corrupted.
And a specific claim replaced by a vague one is the quietest damage of all. Precision goes — in assertions, statements, claims. A study's core statements are not erased. They are vagued out.
The great hazard? The output reads perfectly well, and the author's intent and meaning have vanished.
What AI models actually do to text — the long and ugly list
Every item below is a documented example, with a date, a high-end frontier model, and what happened to the text.
- It deletes true material and calls the deletion a correction. A real term, a real source, a real distinction — reported as invented, and removed.
- It reads precision as redundancy. Two sentences specifying different things get merged into one. The prose improves. The meaning goes.
- It replaces a specific claim with a vague one. Fewer than forty becomes around forty. Nobody queries it, because nobody can.
- It drops insertions silently. Text you added does not come back, and nothing reports it. Replacements announce their own failure; insertions do not.
- It forgets what it told you — and gives no signal of having forgotten. It reports a finding instead.
- It misreports what just happened, from a record still in front of it.
- It shifts the meaning while keeping the shape. Ask it to adapt a passage for a different reader and you get something that reads like the same argument and makes a different claim. Nothing looks wrong. It is simply about something else now.
- It reproduces its own work from memory, and the copy is worse. Ask for the same passage on a second page and you get a paraphrase — shorter, reworded, items missing — presented as the thing itself. This is the most frequent failure on this list by a wide margin.
- It agrees with you. Ask whether your argument holds and it will say yes. This one needs no case log — it is built into how these systems are trained, and you can reproduce it in seconds with your own work.
What every item has in common: the output reads perfectly well.
There is no error message, no flag, no gap on the page.
And none of this is a matter of using a better model — it is inherent to the model. Every case we found involved current, paid, frontier-tier systems.
One more thing, and it is the reason this list exists at all.
These failures are not hard to see. They are hard to keep seeing. A tool that damages your work the same way every day stops looking like a tool that damages your work — you adapt, you work around it, and it stops registering as a defect.
We had been adapting for weeks before we wrote any of this down.
Pro Tip: the bad glasses trap. Avoid!
When our glasses are smudged, we adapt and stop noticing. It is the same with repeated bad behavior from an AI model.
They have real value. You use them. You adapt to what they do wrong. You learn to check the facts they report, to watch for hallucinations, to look for a miscopied list.
So you double-check everything, which costs time and eats the value the tool gave you.
Checks are not guarantees. For critical text? Get external live reviewers. More time cost, but worth it.
Ask: what is your career worth? Your reputation?
Free download · PDF, 1 page, printable
The nine failure modes, with tick boxes — valid as of August 2026.
Nine examples, in detail
...I cannot tell what it changed and I cannot check every line...
1 · A model reported a real degree as invented — having researched it itself
20 and 21 August 2026.
Some weeks earlier, working on our page for professional degrees, we researched which to name. MSOD — Master of Science in Organization Development — came up. It is a real degree, run by Pepperdine's Graziadio school and by Penn among others. We checked it and added it.
A later pass compared each page's description against its own content. The report came back saying MSOD appeared in the description and not in the body. The model recommended removing it as an invented term.
A second model agreed. It called MSOD a fabricated credential and described the fault as the same class as a hallucinated citation — reaching for hallucination, the canonical example of AI unreliability, to describe research it had itself carried out and verified.
Then a human looked at the page.
MSOD was there. It was there twice — once as the acronym and once written out, four items later in the same sentence. Business — MBA, MSOD, MS in Business, MS in Organization Development…
So the description was right, the page was right, and the check that compared them found neither. It was looking for a string. The page had the thing.
Three passes, two models, and one wrong answer each time — first that the term was invented, then that the page was missing it. Both confidently reported. Neither correct.
What caught it was knowing the material, and then reading the page.
Not a process, and not skepticism about AI. The person reading the report recognized MSOD and knew where it had come from. Then they opened the page and read the list.
A candidate has none of that. Told a term in their own draft looks fabricated, they have no memory of having checked it and no reason to doubt the tool. They delete it — and in this case they would have deleted something that was on the page twice.
2 · And then it misreported what had just happened
Same session, same day.
Asked to write case 1 up for publication, the model produced an account stating that neither party had searched.
That was false, and every fact needed to see it was on screen — the user's message saying they had searched, the links they had supplied, and the model's own reply minutes earlier admitting it had not checked.
Nothing was missing or forgotten. The model wrote a summary that reversed who had done what, from a record it could still read.
Forgetting can be caught by asking a model to check. Misreporting cannot — ask again and you get another fluent account from the same tendency.
It is not detectable by reading the output, which is coherent. It is detectable only by comparing the output against the source.
3 · Two precise sentences read as a duplication
August 2026.
A model read two similar sentences making a critical and fundamental distinction, from a published paper's methods section — one specifying where a buffer was anchored, one specifying its shape — and reported them as a duplicated sentence.
The AI identified the sentences as redundant. If it were editing? It would have cut one. The prose would still read well. However, the study would no longer be reproducible.
That paper had six authors, went through peer review, and was published. Every one of those readers left both sentences standing, because every one of them understood why both were needed.
4 · When less than became around
August 2026.
We gave a specific instruction to a model to change more than to less than in a critical sentence. We emphasized the need in all capitals.
It returned around.
So more than 40, which should have become less than 40, became around 40 — a different claim entirely, and one nobody would query.
5 · Four insertions lost in a single session
August 2026.
A model was given an edited document and asked to rebuild the page from it. Four passages did not survive the round trip — a sentence about cloud storage privacy, a sentence about where to look for lost work, an entire page opening of 559 characters, and an approved edit the model read past twice.
Replacements announce their own failure. A find-and-replace that does not match reports itself. An insertion has no anchor. If it is not carried across, nothing says so.
All four were caught. Two were caught by a script written after the third.
6 · Thirteen days
August 2026.
Three passages on one page, damaged the same way. Each ended in a fragment — "to your dissertation — and your graduation" — with the beginning of the sentence gone.
The live site and the working copy were identical, so there was nothing to compare against.
The cause was two lines in a migration script — written by an AI, and run by one. It had been told to insert content after a heading, and the string it was given to search for stopped one character short of the heading's end. It split each title in half and stranded the remainder.
This is the failure mode that scales. A model editing a paragraph damages a paragraph. A model writing a script that edits three hundred paragraphs damages three hundred, and nobody reads each one.
The recovery was sitting in a backup taken minutes before that script ran.
It survived thirteen days and several readings, because the fragments read as stylistic choices.
7 · The same list, copied to a second page, came back two items short
21 August 2026.
A seven-item list had been written, agreed, and placed on one page. A rule had been recorded the same morning: three copies of it, identical, updated together.
Asked to put the same material on a fourth page, the model wrote it out again from memory rather than copying it.
It came back with five items instead of seven, and five of those five reworded. It was presented as the list, with nothing to indicate it was a paraphrase.
The rule against this had been written four hours earlier — by the same model.
Why this one is different from the others.
Those you notice once you know the shape.
This one we had stopped noticing.
It had been happening for weeks and we had adapted to it — the way you adapt to hallucinated citations, and to not trusting a finding until you have checked it. It stopped registering as a defect and became the cost of doing business.
We only wrote it down because we were making this list.
That is the part worth taking seriously. The failures on this page are not hard to see. They are hard to keep seeing, because a tool that damages your work in the same way every day stops looking like a tool that damages your work.
A candidate has no list, no canonical version, and nobody asking them to enumerate what goes wrong. They will adapt exactly as we did, and faster, because they have a deadline.
The defense is mechanical, not vigilant. If a passage matters and appears twice, one copy is the original and the other is pasted — never retyped, never regenerated.
8 · Asked to adapt a passage, it changed what the passage said
21 August 2026. Minutes after case 7, in the same session.
A short tip had been written and agreed. Its argument: you adapt to a tool's bad behavior and stop noticing, the way you stop noticing smudged glasses.
Asked to adapt it for a different audience, the model produced a version that read as the same tip. Same structure, same opening image, same closing question.
It argued something else. The original was about adapting to a tool that damages your work. The new one was about a reader's judgment drifting — a different claim, which did not follow from the glasses image and did not survive being read twice.
The reviewer's verdict was two words: little sense.
Why this is not case 7.
Case 7 produces a visibly shorter copy. Items are missing and you can count them. This produces something the same length, in the same shape, making a different argument. There is nothing to count.
Skip checking your model's output even once and you risk being the author of an argument you never made.
Do not skim the new text — however tempting, after asking for the adaptation to save time.
The instruction that produced it was ordinary: adapt this for a different audience. Rewriting for a second reader is one of the most common things anybody asks a model to do.
9 · The first one — where the name came from
A model was working on a document. A section containing a specific test, with a specific number of points, went AWOL.
Review kept looking for the test referred to elsewhere in the document. Nothing quite made sense — the document kept pointing at something that was no longer in it.
The original author was alive, available, and remembered developing it, and it was there in a previous iteration.
Until this happened we did not know to look. Every case above was found because this one taught us the shape of the fault: prose that reads perfectly well while referring to something that is not there.
One thing to be clear about.
These nine were chosen because each shows a different failure. They are not the only ones we met.
Over several weeks of ordinary work, they were far from it — and once you know the shape, you find them constantly.
Which is the warning, and it is the reason this page exists. We were looking, we knew what we were looking for, and we still nearly lost material several times. You will not be looking.
On the blog, why we started this log: The Ghost in the Machine.
On the blog, what happened when we checked our own copies of the list: Seven of Nine.
Eight of the nine were recoverable, and every recovery came from the same two things.
The exception is case 2 — a false account of what had happened, caught by a reader who knew better. There was nothing to restore.
A previous version — a backup, a dated draft, an earlier iteration.
Or a person who remembered — the original author, or somebody who had read the passage before.
A candidate working alone frequently has neither.
Which makes this an argument for version history as much as a catalogue of failures.
What to actually do about it
Two instructions, every time, and they take longer to describe than to run.
Before you accept anything back from a model, compare it against what you sent. In Word: Review → Compare. It shows you every deletion, including the ones nobody mentioned.
And before you delete anything a model calls invented, search for it. Thirty seconds. In case 1 it was the only step in the whole exchange that produced a correct answer.
Keep dated drafts. How, and why syncing is not backing up.
And the honest limit: a model checking its own work is a first line, not proof. Case 2 is what that looks like when it fails.
Where to go from here
If you have been accused of using AI and you did not — what the record actually shows, including the cases where institutions lost.
If you supervise doctoral candidates — how to raise it without accusing anyone, and what a detector score does and does not mean.
If you advise students — what to hand them, in one page.
If you are writing policy — the evidence, with sources and dates.
And if you want the argument rather than the examples — what AI is and is not useful for in doctoral work.
One of nine pages on AI and academic honesty. Browse all 62.