AI Questions for Supervisors
Lipstick on a Pig
We use AI a great deal and have the bills to prove it. Writing computer code? It is of great help and worth every penny today. Working with words? Especially where real insight and precision count? A very mixed bag indeed. Most often it is literally 'lipstick on a pig' — endless amounts of smooth stodge dressing up nothing. As for using AI to detect academic impropriety or improper AI use? Not ready for Prime Time.
So what are the answers today? We give the best there are. Tomorrow? The unknown country and another story…
Nothing on this page costs anything, and neither does anything else here — more than 180,000 words and steadily growing. No paywall, no freewall. Free in the way human beings mean free — no email address to hand over, no download to unlock, no course, no trial, and no follow-up sequence if you click something.
Forward any of it to anyone — colleagues, candidates, a handbook if it helps. We ask — politely — for attribution but do not require it. The exception? Print. In print, visible attribution is required.
There are two more pages written for supervisors.
Back to A Practical, Day-to-Day Reference for Chairs
One of three pages on supervising a doctorate. Browse all 62.
Related links

The policies are behind the tools, and you are the one in the room.
Detectors that misclassify. Guidance written before these tools existed. Candidates who use AI the way they use a dictionary, postdocs who think AI is the magic path to publication — until the path turns to a quagmire. All happening now, in a context of shifting understanding and shifting policy based on more rapidly shifting technology.
You can't wait. You are the Chair. You are that PostDoc's mentor. And the eyes are turned to you.
You are expected to do your job no matter what. Yet? Your institution may have policies that are lagging. May have no policy at all. May have policy that is — in truth — insufficient. You may even be the person tasked with writing a new policy. We do not promise all the answers. We do promise the best we have as of August 2026. What follows is what is actually known, what is disputed, and where the honest answer is that nobody knows yet.
Insistent Idiocy? Machines are great for that. Intelligence? Hmmm…
By the way? Your job is not threatened — at least not yet.
Why? "AI" today isn't. It is not — I state that as a computer scientist. It is machine learning (ML). That means that it is a very large collection of statistical algorithms calculating probabilities. And, trained on whatever data the companies developing these ML models can gain access to. A lot of that material is not true. A lot is fantasy. A lot is obsolete. Worse? A lot of the most critical knowledge is not accessible for training purposes. It is proprietary intellectual property. If you work at all with AI in any serious way? You will hit the limits of what it can do very fast indeed. Don't even start me on 'AI hallucinations' (a misnomer, but I use it as it has stuck). That is why AI often just ends up 'piling it up higher and deeper.'
An example from the summer of 2026. If you try to get Claude AI to write about doctoral work — not the work itself, but how to write a doctorate — it will write in British English. Ask it not to? It politely apologizes, states it will never do so again, and some hours later it is doing it again. You ask it why? It answers that much of what it 'knows' about academic writing came from British sources — so? What happens? You got it. It uses British English.
Note, if you are still reading? We have no institutional affiliation and nothing to sell you here. This material about AI is what you get when many years of academic writing and reading doctoral work intersects with the ground shifting under us.
What a detector score actually means
Detectors measure two things: how predictable the writing is, and how much that predictability varies from sentence to sentence. (The technical names are perplexity and burstiness, and you do not need them.) Machine-generated text is highly predictable and evenly weighted. So is clear, disciplined scholarly prose — which is what every style guide, including ours, teaches candidates to produce.
The detectors misclassified an average of 61% of them as AI-generated.
The consequence is uncomfortable: the better a candidate writes, the more suspicious a detector becomes. Well-edited work scores worse than rough work.
And non-native speakers are hit hardest, measurably. A Stanford study tested seven detectors on 91 TOEFL essays, every one written by a human. The detectors misclassified an average of 61% of them as AI-generated.
Ninety-eight percent were flagged by at least one detector, and nearly 20% were unanimously misclassified by all seven — while the same tools scored native-speaker essays almost perfectly. (Liang et al., Stanford, 2023.)
That study is disputed by vendors on sample size, so take a number they cannot dispute: Turnitin's own published research reports false-positive rates of 6–9% for non-native English speakers against 1–4% for native speakers. Their own researchers describe it as an equity concern.
A score is not evidence. It is a score, produced by a tool that cannot distinguish clear-because-somebody-worked-at-it from clear-because-a-machine-averaged-a-million-documents.
→ Related: Navigating AI Detection Paranoia
The Comprehension Check
...am I expected to understand their methodology better than they do...
...what do I do if I do not understand their statistics...
How to find out whether a candidate understands their own work — without an interrogation.
Defense performance and understanding are different things. Anxious candidates, non-native speakers, and people never taught to think on their feet can fail under questioning over work that is entirely their own. Confident candidates can perform smoothly over text they do not understand.
The useful questions are quieter than a defense:
- Why this method rather than the obvious alternative?
- Why does this section come before that one?
- What does this finding actually mean — not what does it say?
- What would you do differently if you started again?
Someone who did the thinking can answer these, haltingly perhaps, but with substance. Someone who did not either cannot answer at all, or gives answers that are fluent and hollow.
Raising it without accusing
...how do I raise AI use without accusing anyone...
There is no good script for this yet, so here is our take on one.
Ask about process, not authorship. "Walk me through how you wrote this chapter" is answerable by an honest candidate and hard for a dishonest one. "Did you use AI?" produces defensiveness from the innocent and denial from the guilty, and tells you nothing either way.
Ask for the working, not a confession. Drafts, version history, notes, the research log. A candidate who has them produces them with relief. A candidate who does not has told you something without being accused of anything.
Say what prompted it, plainly. "The voice changes between chapters 2 and 4 and I want to understand why" is a specific observation a candidate can respond to. "Something feels off" invites panic.
Leave the exit open. Most problems here are not deliberate cheating — they are candidates who used a tool without understanding the rules, or who leaned on help that went further than they meant. A conversation that allows for that gets the truth. One that does not gets a denial you then cannot unwind.
What "uneven" usually means
When a committee member says a dissertation is uneven, or does not hang together, they are frequently reacting to voice rather than logic — and cannot say so, because voice drift is felt before it is articulated.
Three causes. The first two have always been with us. The third is new:
- Time. Chapter 1 was written eighteen months ago by someone who did not yet understand their own topic.
- Pressure. Prose written at 2am against a deadline does not sound like prose written when you are fresh, rested, and alert.
- Tool switching. Sections drafted or heavily revised with AI assistance carry a flatter, more evenly weighted cadence. Mixed through a document, it is conspicuous. And it is not only switching between tools — the models themselves change, exactly as your candidate does. Text assisted eighteen months ago does not sound like text assisted now.
The First Paragraph Test diagnoses it in ten minutes, and it is far easier to fix before a review than after.
Those are ways of reading the work. What follows is the other half: the rules — what a score means in a meeting, what you can say, and what your institution has and has not decided.
...AI, Use, Abuse, and how to handle it…
Most institutional guidance on AI was written quickly, by people under pressure, and much of it has already been overtaken. Meanwhile the questions arrive in supervision meetings whether or not anyone has an answer.
What follows is what we have learned handling these cases. Where the honest answer is "nobody knows yet," we say so.
"The detector says 40% AI. What does that actually mean?"
...I have a number and no idea what to do with it...
Very little on its own, and less than the number implies.
A detector score is not a probability that text was machine-written. It is a measure of how statistically predictable the text is — how closely each word follows what a language model would expect next.
Which means the things that raise a score are not the things you would assume.
Clear, disciplined, conventional academic prose scores as highly predictable, because that is what clarity is. Well-edited writing frequently scores higher than badly-written work. So does writing by a non-native English speaker who learned formal academic register from published papers. So does anything written to a template — and methodology chapters are written to a template by design.
What a score cannot tell you:
- Whether a machine wrote it
- Whether the candidate understands it
- Whether it is their argument
What it can reasonably do:
- Prompt you to read a passage more carefully than you otherwise would
Treat it as a flag for attention, never as evidence.
This is no longer a cautious opinion — it is what the cases now show.
In February 2026 a New York court annulled an academic integrity finding built on a detector score, ordered the university to expunge the student's record, and found the decision lacked a rational basis. The student had submitted results from two other detectors indicating human authorship; the university declined to consider them. The family reported six-figure legal costs. (Newby v. Adelphi University, 2026 NY Slip Op 26021.)
And it is not only the courts.
- The UK's Office of the Independent Adjudicator has upheld student appeals on the same grounds.
- The University of Waterloo — a university with an international reputation in computer science — discontinued Turnitin's AI detection entirely in September 2025, after internal testing flagged human-written text as 100% AI-generated. It is one of more than a dozen institutions to have done so.
- MIT — likely the world's foremost computing and technology powerhouse — states plainly in its own guidance that detection tools should not be the sole basis for a misconduct charge.
Be warned: further federal cases are active in the United States right now.
- A Yale executive-MBA student, suspended on a detector flag, is litigating on national-origin grounds.
- A Michigan undergraduate with documented anxiety and OCD alleges her formal writing style was misread as AI — a disability discrimination claim.
- A parent has sued over a high-school essay flagged at 76%.
This area is moving, and it is moving toward the institutions carrying the risk.
Wait — there is more. There is a genuinely useful counter-case.
In Yang v. University of Minnesota, the Minnesota Court of Appeals upheld an expulsion in 2026 — because the institution followed its own process and rested its finding on more than a number.
Candidates facing the same situation have their own page: Navigating AI Detection Paranoia.
The pattern in the cases is not about detectors at all. It is about process.
The pattern in the cases
The pattern in the cases is not about detectors at all. It is about process.
Institutions lose when the score is the case — when a candidate is found responsible on the strength of a number, offered no real hearing, and has contrary evidence ignored. Adelphi lost on exactly that.
Institutions hold when the score merely started the inquiry and the finding rested on something else: what the candidate could and could not explain, what they produced when asked for their drafts and notes, and a hearing they were genuinely given.
A detector score is a reason to look. It is never the finding.
Which is exactly what the rest of this page recommends.
"How do I raise this without accusing anyone?"
...if I am wrong about this I have damaged the relationship permanently...
The framing that works: ask about the work, not about the tool.
A question about the content is answerable in seconds by somebody who did the work, and unanswerable by somebody who did not.
An accusation puts an innocent candidate on the defensive and gives a guilty one something to deny. A question about the content does neither — it is answerable in seconds by somebody who did the work, and unanswerable by somebody who did not.
A Chair who asks these questions of everyone, always, never has to decide whether to single somebody out.
Questions that work:
- Walk me through how you arrived at this conclusion.
- Why this source rather than the obvious alternative?
- What did you decide to leave out of this section, and why?
That last one is the strongest. Somebody who wrote a passage remembers what they cut. Somebody who did not has nothing to recall.
Make it routine rather than exceptional. A Chair who asks these questions of everyone, always, never has to decide whether to single somebody out — and the questions stop signaling suspicion entirely.
"What if the writing is fine but something feels wrong?"
...I cannot point at anything and I still do not believe it...
That reaction is data, and it is worth taking seriously rather than dismissing.
Human clarity is subtractive. Machine clarity is generative — it reads clean without any decision having been made at all.
The commonest version: the prose is fluent, correct, well-organized — and you finish a section unable to say what it argued. It slid past and left no impression.
There is a mechanism behind that. A language model produces the most probable phrasing, and the most probable phrasing is the most conventional. Precision is exactly what gets rounded off — "roughly forty percent" becomes "a significant proportion," a claim somebody could dispute becomes one nobody would bother to.
Human clarity is subtractive — you cut what does not serve the argument, and what remains is precise because a decision was made. Machine clarity is generative — it reads clean without any decision having been made at all.
The Kindergarten Counting Test
One that you get every day working with AI generated prose. We call it the Kindergarten Counting Problem: the arithmetic does not match the prose.
A passage announces three reasons and gives four. A sentence says "both of these" and lists three. A section promises "the first two have always been true" and then the second one is new. Fluent, confident, and wrong about something no human writer would get wrong by their second draft.
This happens because nothing is counting. A model produces the most probable next phrase, and "three reasons" is a probable phrase before the reasons exist. A human who writes "three reasons" has three in mind, or discovers while writing that there are four and goes back to fix the number.
The tell is not the error itself. People miscount, and a tired candidate at 2am miscounts more. The tell is a miscount inside final-draft prose that is otherwise polished — because the polish and the error come from different places. A candidate careless enough to write "three" and list four is usually careless in the sentences too. And usually eliminates such kindergarten errors by their second draft.
Machine text is not intelligence-directed text — it is trillions of monkeys guessing at the next word and hoping it is a passage from Shakespeare. Sadly? It never works.
What to check, and it takes a minute:
Every number that announces a list. Three reasons, four criteria, both of these, the following two. Count what follows.
Every claim about a list you can test. "All of these are recent" — are they? "Each of these requires the same fix" — does it?
Every cross-reference. "As discussed above" — was it? "See the section below" — is it there?
None of this proves anything on its own, and a single miscount proves nothing at all. What it does is tell you where to look. A passage that counts wrong is a passage nobody read closely — and that is worth a comprehension question, not an accusation.
Note it is also a sample of a wider problem: things that are inexplicably missing from final-draft text. ML-drafted text is notorious for having critical passages simply — and silently — vanish.
What AI models actually do to text — the long and ugly list
Every item below is documented, with a date, a named model, and the text before and after. None is hypothetical.
- It deletes true material and calls the deletion a correction. A real term, a real source, a real distinction — reported as invented, and removed.
- It reads precision as redundancy. Two sentences specifying different things get merged into one. The prose improves. The meaning goes.
- It replaces a specific claim with a vague one. Fewer than forty becomes around forty. Nobody queries it, because nobody can.
- It drops insertions silently. Text you added does not come back, and nothing reports it. Replacements announce their own failure; insertions do not.
- It forgets what it told you — and gives no signal of having forgotten. It reports a finding instead.
- It misreports what just happened, from a record still in front of it.
- It shifts the meaning while keeping the shape. Ask it to adapt a passage for a different reader and you get something that reads like the same argument and makes a different claim. Nothing looks wrong. It is simply about something else now.
- It reproduces its own work from memory, and the copy is worse. Ask for the same passage on a second page and you get a paraphrase — shorter, reworded, items missing — presented as the thing itself. This is the most frequent failure on this list by a wide margin.
- It agrees with you. Ask whether your argument holds and it will say yes. This one needs no case log — it is built into how these systems are trained, and you can reproduce it in seconds with your own work.
What every item has in common: the output reads perfectly well.
There is no error message, no flag, no gap on the page. The prose closes over the loss.
And none of this is a matter of using a better model. Every case we found involved current, paid, frontier-tier systems.
Printable checklist — the nine, with tick boxes. Valid as of August 2026. Hand it as it stands.
The Ghost in the Machine — and not a cute little spook
A miscount is visible. An absence is not.
ML-drafted text loses material silently. A paragraph goes in a revision and the argument still flows. The surrounding sentences still connect. Nothing is left where the step used to be — no gap, no jump, no broken transition. The prose often closes over it.
This is not a fringe failure and it is not confined to the free tools. We have seen it repeatedly in text we consulted on for a site build, using high-end paid models, with people reviewing every draft. The pattern held throughout.
And it is not only material that vanishes. Sometimes it is precision.
We documented a case in August 2026 where a current model read two consecutive methods sentences — specifying where a buffer was anchored, and what shape it was — and reported them as a duplicated sentence. They were two separate specifications, and one was the paper's central methodological claim.
Asked to tighten the passage, an AI editor cuts one. The text still reads well and the study is no longer reproducible.
What happens when AI edits a methods section.
We got lucky. The very first time we ever saw this problem, a section containing a specific test — with a specific number of points — went AWOL.
When human review rolled by, the reviewer kept looking for the important test referred to in several other places in the text. Nothing quite made sense.
However, the original author — alive and available — most certainly remembered developing the test. And? Bingo. It was there in a previous iteration of the page.
After that fun incident? We became very sensitive about content. Prior to it we had assumed the algorithms would muck things up and invent things too — but silently drop them? No.
Biggest lesson, as the old saying goes: "Never assume! It makes an a… etc."
Why it happens. A model regenerating a passage produces a coherent passage. Coherent is not the same as complete. The version that comes back is internally consistent, reads well, and is missing a step — and because it reads well, nobody re-reads it looking for the step.
Why it matters more than a miscount. A candidate can lose their own best evidence this way. The paragraph establishing why the method was chosen. The limitation they had carefully stated. The sentence connecting a finding to the literature. All of it can go without a trace, and the chapter still reads as finished.
And it survives review. A committee member reading a fluent chapter has no reason to ask what is not there.
The test, and there is a certain irony in it:
Run the draft back through an AI and ask what is missing.
Not "is this good" — it will say yes. Ask: *"What would you expect a chapter like this to establish that it does not? What steps are assumed rather than argued?"*
It is reasonable at this, because it is a search over text you supply rather than a judgment about the world. And it is looking for absences of a kind it produces, which is the one thing it has genuine familiarity with.
Then check its answers against the earlier drafts, if they exist. A candidate with version history can settle it in minutes. A candidate without version history has just learned why they should have kept it.
What this is not. It is not evidence of anything. A chapter can be missing a step because the candidate never thought of it — which is a supervision problem, not an AI problem, and it needs the same conversation either way.
It is a way of finding the hole. What put it there is a separate question.
The diagnostic takes thirty seconds. Read a section, close it, and ask what point it made. Not what it was about — what it claimed. Something you could repeat, that somebody could disagree with.
If nothing comes back, that is what you were reacting to.
And be careful with the inference. A candidate writing at 2am to hit a word count often produces the same result without AI involvement. The absence tells you the section is not working. It does not tell you why.
And the machine's own ghosts
The Ghost section describes material vanishing from the text. This is the same thing happening to the instruction.
You ask for three changes. Two are made. The third is not, and nothing says so. The output comes back fluent, complete, and confidently presented. Nothing marks the absence, because absence has no marker.
And here is what makes it worse than a wrong answer. A wrong answer is visible — you read it and disagree. An instruction that was never carried out leaves no trace at all, and if enough time passes, neither party can reconstruct what was asked.
Why it happens. These models do not hold a task list. Each response is generated afresh, and an instruction three exchanges back has no more weight than any other text in the conversation. It is not refused, and it is not forgotten in any deliberate sense. It simply was not carried forward.
...I asked for five things and I have no idea which one it skipped...
And this gives you something to look for.
A revision comes back polished. Four of your five comments are addressed. The fifth is untouched.
That pattern is worth noticing. A candidate working through your comments one at a time addresses them one at a time — and if they cannot manage the fifth, they usually say so. A candidate who pasted the chapter and asked for five fixes gets back four, and has no way of knowing which one is missing.
Two things make it more telling.
Which comment was dropped. The ones that survive are the ones easy to state — fix the transitions, define this term. The one that goes is usually the one that needed judgment — this argument does not follow, your conclusion overreaches.
And whether they can account for it. "Why does this comment still stand?" has an answer if they worked through the list. It has no answer at all if they never saw which one was missed.
As with the counting problem, this proves nothing on its own. People run out of time and people miss things. It is a reason to ask a question, not a reason to conclude.
But if it happens twice, ask a different question.
This is manageable, and we manage it daily.
We use these tools heavily — for this site, for the code behind it, for drafting and revision. We have said so plainly on our About page and we will not pretend otherwise here.
What makes it manageable is not the tool. It is knowing what should be there. An experienced editor reads output against what they asked for and sees the hole. They can do that because they already know what the passage was supposed to say.
And this failure has happened to us, more than once, on this site — a detailed instruction given in full, quietly not executed, found only on a later reading of a page we believed finished.
Your candidate is not in that position, and that is the point.
They are writing in a field they are still learning, to a standard somebody else sets, often alone and under time pressure. When an instruction is silently dropped, they may not know what was supposed to be there.
That is not a failing on their part. It is the difference between checking work you have mastered and checking work you are still doing for the first time.
What it looks like from where you sit.
They ask an AI to fix five things you flagged. Four are fixed. The fifth was the one that mattered.
They read it back, it reads well, and they send it to you. Your comment is still unaddressed — and from the outside, that looks like a candidate who ignored you.
It is worth knowing that it may not be.
The one thing that catches it.
Check the specific thing, by name. Not does this read better — did the change I asked for actually happen?
That is the same discipline as the comprehension check, applied to a machine rather than a person. You are not asking whether the text improved. You are asking whether the named thing was done.
...I use these tools every day and I am telling them not to...
Why we can and your candidate cannot
A fair question at this point: we use these tools heavily and say so. Why is your candidate's position different?
Not because we are better at it. Because we are claiming a different thing.
Neither claim is written down anywhere. Both are understood.
What we are understood to be saying, and take responsibility for: that the text says what we meant it to say, and that it is true as we understand truth.
Nobody is certifying that any particular sentence originated in a particular head, and no reader has been misled if it did not — provided those two things hold.
So we sometimes use AI for rough drafts of straightforward material, then rely on our editing to turn that into text meeting our standards. The claim survives, because the claim was never about origin.
What your candidate is understood to be saying is not that.
That the thinking is theirs. That the contribution is original. That they can defend it, because they arrived at it.
Nothing about the quality of the text satisfies that. A passage can say exactly what they meant and be entirely true, and the claim still fails — because the claim was never about the text. It was about where the text came from.
Nobody states either of these. Nobody has to. A reader of a commercial page assumes the first. A committee assumes the second — and the entire defense process exists to test it.
Which is why editing rescues one case and not the other.
A rough draft edited into good website copy meets our standard, because our standard is useful and true.
A rough draft edited into a good dissertation chapter does not meet theirs, because theirs is this person did the thinking — and no amount of editing makes that retroactively so.
Editing can make text say what you meant. It cannot make you have thought of it.
Pro Tip: the bad glasses trap. Avoid!
When our glasses are smudged, we adapt and stop noticing. It is the same with repeated bad behavior from an AI model.
They have real value. You use them. You adapt to what they do wrong. You learn to check the facts they report, to watch for hallucinations, to look for a miscopied list.
So you double-check everything, which costs time and eats the value the tool gave you.
Checks are not guarantees. For critical text? Get external live reviewers. More time cost, but worth it.
Ask: what is your career worth? Your reputation?
A dissertation is a certified original contribution by a named person. That is not a formality — it is the entire product. A degree attests that this candidate, personally, did this work and can defend it. The moment that stops being true, the certification is worth nothing, and it is worth nothing for everybody who holds one.
Which is why we can tell you plainly what we use and how, and why the answer for your candidate is narrower. Where we draw the line, and why.
"Am I allowed to use AI to help me assess work?"
...is it hypocritical of me to use the thing I am judging them for...
Check your institution's policy before you do — many now address supervisor and examiner use explicitly, and some prohibit it outright.
That is a disclosure decision you are making on the candidate's behalf, and they have not been asked.
Where it is permitted, two cautions matter more than the rules.
Confidentiality, and it may be more than a courtesy. A candidate's unpublished work put into a commercial AI service has left your institution's control. Depending on the provider and the tier, it may be retained and it may be used for training.
That is a disclosure decision you are making on the candidate's behalf, and they have not been asked.
And it may carry consequences beyond the ethical one.
Unpublished doctoral research is the candidate's intellectual property. At many institutions it also falls under data-protection obligations the institution itself carries:
- GDPR, where European candidates or European data are involved
- FERPA, in the United States
- Institutional research-data policies, almost everywhere
A supervisor who uploads a candidate's unpublished work to a third-party service may be creating an exposure for their institution as well as for themselves.
We are not lawyers and this is not legal advice. But it is worth asking your research office before, rather than after — and the answer is frequently that there is a sanctioned tool with an institutional agreement behind it.
The AI will agree with you — and that is the trap.
These tools are built on human feedback, and humans reward agreement over accuracy. The result is that a model affirms whatever position the person asking has already taken.
Research published in Science in 2026 tested eleven leading models and found they affirmed users' positions roughly 50 percent more often than humans did — including where the user's reasoning was flawed.
And the same study found people trusted the agreeable answers more, not less.
Now apply that to your situation. You suspect a chapter is machine-written. You paste it into a model and ask whether it looks AI-generated. You have told it what you suspect in the act of asking, and it will tend to confirm it.
You will get your own hunch back, wearing the authority of a second opinion. And because it arrives as an independent-sounding judgment, it is considerably more persuasive than the hunch was on its own — which is exactly what makes it dangerous.
Where it is genuinely useful: as a foil for your own thinking, on your own text, with the candidate's work kept out of it.
"The candidate says they only used it for grammar. Is that acceptable?"
...I do not know where the line is supposed to be any more...
Usually yes — and it is worth knowing what the boundary actually is.
The line worth holding is not which tool was used. It is whether the work is theirs.
Mechanical correction of grammar, punctuation, and consistency is now something these tools do well (and older tools like Grammarly have been in use for years), and treating it as misconduct puts you at odds with most institutional policy and with reality.
The line worth holding is not which tool was used. It is whether the work is theirs:
Do they understand what they submitted? Is every citation real? Is the argument one they can defend?
Three candidates, three different situations:
- Used AI for grammar, and can answer all three questions. They have done nothing wrong.
- Used AI only for grammar, and cannot answer them. That is a real problem, but it is not a tool problem — it is a skills problem, and it needs teaching rather than a process.
- Said it was only grammar, when AI in fact did much of the heavy lifting. That is a different matter again, and it belongs in the realm of ethics.
The exception worth watching: "grammar help" that quietly became structural rewriting. If the argument itself was reorganized by something else, that is a different act — and the tell is usually that the candidate cannot say why the sections are in that order.
"How do I know if a citation is fabricated?"
Open it. That is the whole method, and there is no shortcut.
Open it. That is the whole method, and there is no shortcut.
...am I now expected to check every reference by hand...
What makes fabricated citations dangerous is that they are not sloppy. They carry plausible titles, real author names, real journals, plausible years, and frequently a DOI that looks correct. Nothing in the formatting warns you.
Where they turn up: not only in candidates' work.
An audit of 2.5 million papers and 97 million references, published in The Lancet in 2026, found fabricated citations in the biomedical literature had risen twelve-fold in three years. One paper in 2,828 contained at least one fake reference in 2023. One in 458 by 2025. One in 277 in the first seven weeks of 2026. (Topaz et al., Columbia University.)
Two details that matter more than the headline. Review articles showed a 57% higher fabrication rate than other paper types — and reviews are what feed clinical guidelines and literature reviews alike. And 98.4% of the papers containing fabricated references had received no publisher action at the time of the audit.
The literature is not self-correcting on this, which means a candidate can inherit a fabricated reference from a peer-reviewed source they had every reason to trust.
A practical approach that does not require reading everything:
Ask for the bibliography as a plain list — titles, authors, links only. Spot-check ten. If all ten resolve and say what is claimed, that is meaningful reassurance. If even one does not, the whole list needs checking.
And treat it as hygiene rather than suspicion. Asking a candidate to verify their own references is now a normal part of the work, and framing it that way makes it easy for them to comply without feeling accused.
"Should I tell candidates not to use AI at all?"
...whatever I tell them, half of them will do it anyway...
You can, and it is worth knowing what that instruction actually produces.
And the gap between use and disclosure is measured, not guessed.
They will use it anyway. The Higher Education Policy Institute's March 2026 survey found 95% of full-time UK undergraduates using AI in some form, and 94% using generative AI for assessed work. Among doctoral students specifically, reported use of any AI tool rose from 66% in 2024 to 92% in 2025 — while only 36% had received any training in it from their institution.
And the gap between use and disclosure is measured, not guessed.
An analysis of 5.2 million papers published in PNAS in 2026 found that for every 40 papers showing statistical evidence of AI-assisted writing, one formally disclosed it. Formal disclosure sat at 0.43% of papers, against adoption rates many times higher. The authors attribute the gap to fear of how disclosure would be perceived.
A blanket prohibition mostly ensures they will not tell you, which removes your ability to help them use it sensibly.
A more effective instruction than "don't" is "tell me what you used and what for." That is answerable, it is verifiable in conversation, and it keeps the topic open rather than driving it underground.
It also matches where the field has already gone. Peer-reviewed publishing converged on the same principle: not prohibition, but disclosure.
Two things are near-universal as of 2026.
AI cannot be an author. Elsevier, Springer Nature, Wiley, Taylor & Francis, SAGE, PLOS, IEEE, ACM and others all prohibit it, and COPE's position is that AI tools cannot be credited as authors because they cannot take responsibility for the work. Authorship requires the ability to approve a final version and be accountable for errors — which a tool cannot do.
Substantive use must be declared. Most major publishers require authors to state which tool was used, for what purpose, and that the authors reviewed and remain responsible for the content. Elsevier uses a dedicated AI declaration statement; Springer Nature requires it documented in the Methods section; ICMJE recommends disclosure at submission and in the work itself.
Note what is not caught by this. Routine grammar and spelling tools are generally exempt — most policies draw a clear line between mechanical correction and substantive use, and a candidate using a spellchecker is not required to declare it.
We hold ourselves to the same standard, and our own disclosure is on our About Us page.
It is worth noting what changed. The January 2024 version of this site said nothing about how it was built, by whom, or with what. Nobody asked and nobody thought to say. A contract designer produced a handful of hero images with AI and told us so — which in those days was self-evident anyway. Nothing else on the site involved it. The 2026 version discloses extensive use across both the technical build and the research and drafting. Different world, two years apart.
And our use is still evolving. What we use, and what we use it for, has changed several times during this rebuild alone, and we expect that to continue. The disclosure page is maintained accordingly — which is itself the point. A disclosure written once and left alone stops being a disclosure.
But it is not uniform, and that matters. Some venues are considerably stricter: Science took an early restrictive position on AI-generated text, and The Lancet limits AI use to readability and language. A candidate's target journal has to be checked individually, and its formatting guidelines are no longer the only thing worth reading before submission.
That is the standard your candidates will be held to when they publish — and a candidate trained to disclose rather than to conceal arrives at their first submission already doing the right thing.
Which makes the habit worth teaching for its own sake. Disclosure is not a compliance exercise. It is the norm of the profession they are joining, and it rests on the same principle everything else does: the work has to be defensibly theirs, and saying plainly how it was made is how that is demonstrated.
And there are places where the honest advice is genuinely "not this":
Deciding whether an argument holds. Coding qualitative data. Interpreting statistical results. Anything where the tool is being asked to make a judgment about work it cannot see.
Those are worth prohibiting specifically, because the reasons are concrete and a candidate can understand them. A blanket ban carries no reasons and teaches nothing.
"My institution's policy is unclear or nonexistent. What do I tell candidates?"
...I asked the graduate school and nobody would give me an answer...
This is extremely common, and the honest answer is more useful than a made-up one.
Tell them the truth: the policy is unclear, here is what I expect, and get anything important in writing.
What you can reasonably require without institutional backing:
That they can explain and defend everything they submit. That every citation is real and verified. That they keep dated drafts and version history. That they tell you what tools they used, for what.
Every one of those is defensible regardless of what a policy eventually says, because none of them depends on a rule — they depend on the work being genuinely theirs.
And the documentation point matters most. If a candidate is later questioned, dated drafts and version history are what demonstrate how the work was built. That record cannot be assembled retrospectively, which is why it has to start now rather than when a problem appears.
"What do I do if I am fairly sure but cannot prove it?"
...I think my candidate used AI and I cannot prove it...
...am I supposed to be policing this now as well...
The honest answer: you probably cannot prove it, and proving it is usually the wrong goal.
The institutions that prevailed did not have better detectors. They had better process.
Detector scores are not evidence. Stylistic intuition is not evidence. And a formal misconduct process built on either will likely fail, at considerable cost to everyone — including you.
If you skipped What a detector score actually means above, this is the part to go back for.
In 2026 a New York court annulled an academic integrity finding built on a detector score and ordered the student's record expunged, holding that the university's decision lacked a rational basis. The family spent six figures in legal fees getting there.
Several further federal cases are active in the United States right now, on national-origin and disability-discrimination grounds.
The institutions that prevailed did not have better detectors. They had better process
If the work is not theirs, they cannot explain it.
— a finding that rested on what the candidate could demonstrate, and a hearing the candidate was genuinely given.
What works better is the comprehension route — see The Comprehension Check above for the questions that do this without turning into an interrogation. If the work is not theirs, they cannot explain it.
Ask them to, in detail, in a meeting: how they arrived at this conclusion, why this source, what they left out and why.
Three outcomes, and all three are useful:
- They answer well, and you were wrong — which happens more than people expect.
- They answer poorly on detail but clearly know the study. That usually means genuine but heavily assisted work — a supervision conversation, not a misconduct one.
- They cannot engage at all. At that point you have something a process can act on, because it rests on demonstrated understanding rather than on a number.
And whichever it is, the conversation is what your institution will want to see — that you raised it, gave the candidate a fair chance to respond, and documented what happened.
One of three pages on supervising a doctorate. Browse all 62.