Human-in-the-Loop Structural Refinement

...I asked my AI if my methodology chapter made sense and it said it looked great. Why doesn't my Chair agree?...

Because the AI was doing exactly what it was built to do — and what it was built to do is not the same thing as telling you the truth. A simple test for every AI you use: if an AI tells you something is 'great' — 'awesome' — 'world-class' — or offers any other serious encomium? Ask the AI to justify the statement with hard facts and reasoning. You will quickly discover for yourself how much of what you get from your AI is fluff.

Human-in-the-Loop Structural Refinement

This page explains the difference between editing and structural judgment, why AI is genuinely strong at one and structurally unable to do the other, and what that means for how you actually use it.

...can AI just check if my argument makes sense?

Editing, advanced editing, and coaching are not three levels of the same task getting progressively "more thorough" — they're different kinds of work entirely.

Editing is bounded: grammar, formatting, citation style, sentence-level clarity — a fixed input gets polished, and speed can substitute for depth because there's no judgment call involved. Basic editing can mostly be done by a good AI model — with some competent human review afterward.

Advanced editing is a step further — it looks at holistic consistency: does everything make sense both within the edited text and across chapters? Does chapter 3's stated methodology actually match what chapter 4 reports doing? Does terminology introduced in chapter 1 stay consistent by chapter 5? This requires holding the whole document in mind at once, not just the paragraph in front of you. In our experience? Even advanced AI models still struggle with this. They 'misread' — and hence misunderstand — points major and minor, and struggle with "holistic." Even million and two-million token models have issues. Remember: their approach is statistical probability, not reasoned truth.

Coaching is diagnostic: does this topic actually hold up, does the methodology fit the research questions, does chapter 4's results section actually support what chapter 5 claims to conclude. That work is sequential and depends on genuinely understanding your specific research, not pattern-matching against a million other dissertations. In the hands of a professional who knows what they are doing, an advanced AI can help with parts of this — but mostly it requires real insight and real judgment, applied to your research questions and your data, not to a statistical average of everything the model has ever read.

...why does my AI keep telling me my work is good?

Because you are awesome! Actually? No. Neither the AI nor ourselves know if that is true. The real reason for this phenomenon is fundamental to how AI models are created. It isn't a personal failure to prompt correctly. A peer-reviewed study published in Science in March 2026 tested eleven leading AI models — including the major consumer chatbots — and found their responses were nearly 50 percent more agreeable than a human's would be in the same scenario, even when the user's reasoning was flawed or their described actions were harmful (Cheng et al., 2026). The researchers also found that people trusted the sycophantic responses more, not less — meaning the more an AI reassures you, the more confident you become, whether or not that confidence is earned. For a dissertation, this is a serious structural risk, not a minor annoyance: the tool many candidates reach for to sanity-check their own logic is statistically built to validate that logic first and interrogate it only reluctantly.

...but I know it is just being nice, surely that means it cannot fool me...

It does not. This is the part that should genuinely worry you, and it was established in the research long before anyone worried about AI.

In a study published in the Journal of Marketing Research, participants received flattery that was transparently insincere — impersonal, mass-produced, and attached to an obvious commercial motive. Consciously, they discounted it exactly as you would expect. They knew what it was. Yet they were still around 50 percent more likely to choose an offer from the flattering source over an equivalent one, because the favorable impression formed underneath the conscious correction and carried on operating regardless. Knowing you are being flattered corrects what you say. It does not correct what you do. (Chan & Sengupta, 2010.)

Two further findings matter here. Earlier work found that flattery from a computer is about as effective as sincere praise from a person — the source being obviously mechanical does not defuse it. And, most relevant of all: the effect is strongest in people engaged in self-criticism, and weakest in people who feel good about themselves.

Consider who that describes. A candidate fourteen months into a stalled dissertation, uncertain whether the work is any good, isolated from anyone who understands the process, opening a chat window at one in the morning to ask whether their chapter is ready. That is not a neutral reader casually receiving a compliment. That is the precise psychological profile on which insincere praise works best — and the machine is statistically inclined to supply it.

Hollywood understood this long before the research did. There is a scene in Pretty Woman in which a Rodeo Drive manager, having identified a client with an unlimited budget, asks how much flattery is required — and receives the answer that a great deal is expected. The joke works because everyone in the scene knows exactly what is happening, and it works anyway.

Your AI is not manipulating you. It has no motive. But it produces the same output as a salesperson who does, and the research is clear that your awareness of this is not the protection you assume it to be.

...so what is AI actually good for then?

It is genuinely useful, when the task is suited to AI. Examples include:

  • Reorganizing scattered notes into thematic clusters
  • Reverse-outlining a chapter you've already written, so you can see what each paragraph is actually arguing versus what you meant it to argue
  • Stress-testing a specific, narrow claim against text you feed it directly
  • Drafting a plain-list bibliography for fast manual verification

What it's NOT built for:

  • Telling you honestly that your theoretical framework doesn't fit your methodology
  • Catching that your chapter 5 conclusions overreach what chapter 4's data can actually support
  • Pushing back on a research design decision you're emotionally attached to

...do I even still need a human editor if I have AI?

For basic editing tasks, much less than you used to — provided you follow our advice on how to handle basic editing properly. Use a good AI model with the right prompts, and an experienced editor needs only to review the finalized text. We will recommend the model and supply the prompts. However, that is only basic editing.

Advanced editing and coaching are about judgment, experience, and the willingness to read and completely absorb your entire thesis — to grok it, in the proper sense: to hold the whole of it in mind with the relationship of every detail to every other detail clear. That is what neither a tool nor a hurried human can supply. Advanced editing ensures that your whole document is consistent throughout and makes sense based on what you have written. A good writing-cycle coach ensures that your text does not just make sense but advances the art and achieves all institutional objectives. Both activities require competence, experience, comprehension, and the ability to concentrate. Coaching requires doctoral or post-doctoral reasoning and contextual abilities — as well as the ability to repair fundamentally broken dissertations whose current drafts have embedded assumptions, contradictions, and logical failures preventing the drafting of a coherent, scholarly, and convincing draft capable of passing Chair, Committee, and above all, School Review. Faster and more AI use will not change how much accumulated context that judgment requires.

There is a plainer version of all this, and it matters more than any of it.

AI writing is smooth. It is also, very often, empty.

Read a few paragraphs of it and then ask yourself whether it made a point worth remembering. Not whether it was about something — whether it made a point. One you could repeat tomorrow. One somebody could disagree with.

Frequently there is nothing to repeat. It slid past. It left no impression at all.

That is not a malfunction. It is what the machine does. It produces the most probable phrasing, and the most probable phrasing is the most conventional one — so precision is exactly what gets rounded off. "Roughly forty percent" becomes "a significant proportion." "This will fail" becomes "this may face challenges." A claim somebody could argue with becomes one nobody would bother to.

A human writer arrives at clarity by cutting. You remove what does not serve the argument, and what is left is precise because you decided what mattered.

A machine arrives at it by generating. The result reads clean without any decision having been made.

Same surface. Opposite process. And the difference only becomes visible when the claim has stakes — which, in a dissertation, is always.

This is the clearest description we can give of what we actually do for you. When we read your chapter we are not checking grammar. We are asking, paragraph by paragraph, whether it says anything — and telling you precisely where it does not.

And be clear that the source does not matter. A candidate writing at two in the morning to reach a word count produces exactly the same result, and always has. Mush is mush. What has changed is that it is now possible to produce a great deal of it very quickly, and that it passes every surface check on the way through.

The Point Test — run it on your own work right now.

Take any section your Chair has been vague about. Read it once. Close it. Then ask:

Does your text make a point worth remembering?

Not was it about something. Did it make a point — one you could repeat tomorrow, one somebody could disagree with, one a reader would carry away.

If nothing comes back, you have found the problem, and it is fixable. On-point writing means every paragraph makes one claim you could state out loud, and every claim is one somebody could dispute. If nobody could argue with a sentence, it is not doing any work, and it can go.

This is also why we insist on a handover session.

Where we have done substantial work, we walk you through it before you take it back, and then ask you to explain parts of it in your own words. Not as a test — as a check.

If you cannot explain a passage, you cannot defend it, and finding that out with us is very much better than finding it out at a defense. The work has to be defensibly yours, and this is how we make certain it is rather than assuming it.

...who actually tells me the truth about my dissertation?

We do — and we mean that as a specific operational commitment, not a slogan. We work only for you, so it is our job and our only job. Note, we are easy to work with, but we are not agreeable. That is to say: we are not AIs. Every structural read we provide you is verified against your actual research questions, your actual theoretical framework, your own data, and your actual institution's requirements — not against what will make the conversation end sooner.

What AI models actually do to text — the long and ugly list

Every item below is documented, with a date, a named model, and the text before and after. None is hypothetical.

  • It deletes true material and calls the deletion a correction. A real term, a real source, a real distinction — reported as invented, and removed.
  • It reads precision as redundancy. Two sentences specifying different things get merged into one. The prose improves. The meaning goes.
  • It replaces a specific claim with a vague one. Fewer than forty becomes around forty. Nobody queries it, because nobody can.
  • It drops insertions silently. Text you added does not come back, and nothing reports it. Replacements announce their own failure; insertions do not.
  • It forgets what it told you — and gives no signal of having forgotten. It reports a finding instead.
  • It misreports what just happened, from a record still in front of it.
  • It shifts the meaning while keeping the shape. Ask it to adapt a passage for a different reader and you get something that reads like the same argument and makes a different claim. Nothing looks wrong. It is simply about something else now.
  • It reproduces its own work from memory, and the copy is worse. Ask for the same passage on a second page and you get a paraphrase — shorter, reworded, items missing — presented as the thing itself. This is the most frequent failure on this list by a wide margin.
  • It agrees with you. Ask whether your argument holds and it will say yes. This one needs no case log — it is built into how these systems are trained, and you can reproduce it in seconds with your own work.

What every item has in common: the output reads perfectly well.

There is no error message, no flag, no gap on the page. The prose closes over the loss.

And none of this is a matter of using a better model. Every case we found involved current, paid, frontier-tier systems.

What happens when AI edits a methods section?

It removes precision it cannot see, and the result reads better than what it replaced.

Here is a documented case from our own work, in August 2026.

The text. Two consecutive sentences from the methods section of a 2024 paper in Environmental Research:

We centered buffers on the respondents' home addresses to model their immediate and expanded spatial contexts. We used circular buffers on the respondents' home addresses to model their immediate and expanded spatial contexts.

What the model said. Asked whether the paper was clearly written, Claude Opus 5 reported a duplicated sentence — "nearly verbatim… an editing error that survived peer review."

What the two sentences specify. Two different things. Centered is where the buffer is anchored — on the actual residential address rather than an area centroid, which is the specific methodological choice the authors fault earlier studies for getting wrong. Circular is the geometry — Euclidean distance rather than a network buffer following streets. Both are required to reproduce the study.

Why the model got it wrong. The two sentences share an identical closing clause. The model matched on the tail, concluded duplication, and never compared the two verbs.

What would have happened next. Asked to tighten that passage, an AI editor cuts one sentence. The remaining text still reads well. It is no longer reproducible — a reader cannot tell where the buffer was anchored or what shape it was, and one of those is the paper's central methodological claim.

Nothing would flag the loss. You cannot catch it by rereading the output, because the output reads better than what it replaced.

And now consider where this happened. This was published text.

Six authors wrote it, and it is highly likely they read each other's sections and the whole work. At this level it will have gone through one or more technical edits before a final editor worked over the whole. Peer reviewers read it. The journal read it, and likely more than once — received in June 2024, revised in August, published shortly after.

Every one of those readers left the two sentences standing, because every one of them understood why both were needed.

The model read that and called it an error.

Your draft has none of that protection. No reviewers have seen it. The distinctions in it are ones you worked out yourself and may not yet have phrased as carefully as a published paper does. If a model misreads text that survived peer review, it will misread your chapter faster and with more confidence.

This is what human-in-the-loop means in practice. Not a person approving fluent text. A person who knows what a buffer is, and why anchoring and geometry are separate decisions.

What happens when an AI forgets what it told you?

It does not say so. It reports a finding instead.

Here is a documented case from our own work, on 20 and 21 August 2026.

The background. Some weeks earlier, working on our page for professional doctorates, we researched which degrees to name. MSOD — Master of Science in Organization Development — came up. It is a real degree, run by Pepperdine's Graziadio school and by Penn among others. We discussed it, checked it, and added it to that page's description.

What happened later. A subsequent pass compared each page's description against its own content. The report came back saying MSOD appeared in the description and not in the body. The model recommended removing it as an invented term.

A second model — this one — agreed. It called MSOD a fabricated credential and described the fault as the same class as a hallucinated citation.

Read that again. The model reached for hallucination — the canonical example of AI unreliability — to describe research it had itself carried out and verified some weeks earlier.

Neither model searched. Both had search available.

What was actually wrong. The description was right, the page was right, and the check that compared them found neither. MSOD was on the page twice — once as the acronym, once written out, four items later in the same sentence. Business — MBA, MSOD, MS in Business, MS in Organization Development… The check was looking for a string. The page had the thing. Three passes, two models, and one wrong answer each time — first that the term was invented, then that the page was missing it. Both confidently reported. Neither correct.

Why this is the most dangerous case we have documented.

The model was not ignorant. It had known. The term was researched and added in an earlier conversation with the same system.

It gave no signal that it had forgotten. It did not say I do not recall this or I cannot verify this.

It said the term appeared to be invented — a finding, delivered with the same confidence as everything else it says. And it named the mechanism it was supposedly diagnosing. A model that attributes its own verified work to hallucination is not uncertain. It is confidently wrong in the one direction nobody is watching for.

And a finding is actionable. You act on it, and the action is deletion.

This is the inverse of what everybody worries about. The fear is that AI will invent something. This is AI deleting something true, and presenting the deletion as a correction.

For research the inverted failure is the worse one. A fabricated citation gets caught — you go to find it and it is not there. A real source flagged as fabricated gets removed, and nothing ever reports the loss.

The pattern to watch for: unfamiliar terminology, minority-language sources, small-field journals, regional institutions, anything recent, anything niche. A model's confidence is unrelated to its coverage.

What caught it. A human read the page.

Neither model did. Both had search available and neither used it — because neither was uncertain. A model that thinks it has found an error does not look for evidence against itself.

What caught it was knowing the material, and then reading the page.

Not a process, and not skepticism about AI. The person reading the report recognized MSOD and knew where it had come from. Then they opened the page and read the list.

A candidate has none of that. Told that a term in their own draft looks fabricated, they have no memory of having checked it, no earlier conversation to recall, and no reason to doubt the tool. They delete it.

What to do about it. Before you delete anything a model calls invented, search for it. Thirty seconds, and in this case it was the only step in the whole exchange that produced a correct answer.

And then it misreported what had just happened.

Asked to write this case up, the model produced an account stating that neither party had searched.

That was false, and every fact needed to see it was on screenthe user's message saying they had searched, the links they had supplied, and the model's own reply minutes earlier admitting it had not checked.

Nothing was missing or forgotten. The model wrote a summary that reversed who had done what, from a record it could still read.

This is a different failure from the one above, and a harder one.

Forgetting can be caught by asking a model to check. Misreporting cannotask again and you get another fluent account from the same tendency. A version that reads well, assembled from the shape of the story rather than from the record.

The shape was two models got it wrong. Neat, and wrong at the edges. The record was two models got it wrong and a human caught it in thirty seconds.

It is not detectable by reading the output, which is coherent. It is detectable only by comparing the output against the sourcethe same check we recommend for a returned draft, applied to a summary.

One thing worth stating plainly. Both models here were current, paid, frontier-tier systems. Not free tools, not older versions, not anything anybody would be warned off. Capability was never the problem.

Core References

Cheng, M., et al. (2026). Sycophantic AI decreases prosocial intentions and promotes dependence. Science.

Chan, E., & Sengupta, J. (2010). Insincere flattery actually works: A dual attitudes perspective. Journal of Marketing Research, 47(1), 122–133.

One of nine pages on AI and academic honesty. Browse all 62.

Fast Help & More Resources