A claim needs a source whenever a reader could ask "says who?" — and AI drafts fail that test in three specific ways. Missing attribution, unverifiable attribution, and misattached attribution each need a different repair. The ForgeRank published dataset (https://forgerankai.com/static/data/ai-content-quality-29-drafts.json) scored 29 drafts and logged 2,839 claims; 51 were flagged as needing a source, and 38 of those 51 carried none. That is a 75 percent failure rate on the claims that mattered most.

What you need before you start

  • The draft as plain text, with paragraph numbers you can reference
  • A browser with two tabs open: one for the source, one for the draft
  • The five-source list from Google Search Central's spam policies page, so you know what counts as scaled content abuse
  • Thirty to sixty minutes per 1,500-word draft; source checking runs slower than reading
  • No special software. A spreadsheet with four columns — claim, source named, source verified, supports sentence — handles a 2,000-word draft

You do not need to verify every sentence. You need to verify every sentence a reader could challenge.

Step 1: Pull every claim that needs a source into one list

Read the draft once and copy out every sentence containing a number, a date, a named organisation, a study reference, or a causal assertion. The ForgeRank methodology page notes that source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap — which means the tooling stops counting before the draft stops claiming. Your list will run longer than you expect. A 1,200-word draft with eight statistics and four "research shows" constructions produces roughly a dozen rows.

Sort the rows into three buckets as you go: no source at all, source named but unverifiable, and source named and real but not supporting the sentence. Those three buckets get different fixes in the steps below.

Step 2: Repair the claims with no attribution at all

For each unsourced claim, decide whether it survives without one. A mechanism explanation — how a process works, in what order, at what cost — carries itself. A statistic does not. When the claim is a number, either attach a source you have personally opened or cut the number and keep the insight in words. "Roughly 61 percent of non-native English essays were flagged as AI-generated" needs Weixin Liang et al., arXiv 2304.02819, published in Patterns, attached inline, because the number is the whole point of the sentence. The same sentence without the number — "AI detectors misclassify non-native English writing at high rates" — stays true and needs nothing. Google Search Central's March 2024 core update, which reported a 40 percent reduction in low-quality unoriginal content, gives you the reason to bother: unattributed filler is the category that update targeted.

Step 3: Test whether a named source is actually checkable

A named organisation is enough for attribution even with no link. "The National Sleep Foundation recommends seven to nine hours" passes, because a reader can find the foundation. "Studies show" fails, because no reader can find "studies." Run each named source through one question: could a stranger locate this in under two minutes using only what the sentence gives them? Organisation names, report titles, author names, and publication years all pass. Vague collective nouns do not.

The trap here is a source that sounds specific and resolves to nothing. "A 2024 industry survey found 68 percent of marketers..." names a year and a population but no publisher, which makes it unverifiable and functionally invented. Rewrite it as mechanism or drop it.

Step 4: Check whether the source supports the sentence it sits on

This is the failure mode that survives every other check. The source is real, the link works, the number matches — and the sentence claims something the source never said. Open the source and read the passage, not the abstract. A study measuring detection rates on TOEFL essays tells you about false positives in academic writing; it does not tell you that AI detectors fail in workplace settings. Attaching Liang et al. to the second claim misuses a real citation, which is harder to catch than a missing one and more damaging when a reader catches it.

Fix by narrowing the sentence to what the source measured, or by finding the source that measured what you claimed. Never widen the source to fit the sentence.

Step 5: Rewrite abstract claims into checkable ones

Two worked rewrites from the same draft, quoted as they appeared:

Before: "Sourcing matters more than most writers think."

After: "A claim needs a source when a reader could ask 'says who?'"

Before: "Quality varies across AI drafts."

After: "Four of the 29 drafts scored 7.2, and each one carried the same 7.0/7.0/8.0/7.0 dimension profile."

The second pair matters because the abstraction hid the finding. Tabulating the four dimension tiers behind every row in the ForgeRank published dataset shows each step in the distribution is one dimension moving. All ten 4.0 drafts scored 4.0 across all four dimensions. All ten 4.8 drafts scored 4.0/4.0/7.0/4.0. The three 7.0 drafts were flat at 7.0. The four 7.2 drafts sat at 7.0/7.0/8.0/7.0. The single 7.8 draft was 8.0/7.0/9.0/7.0, and the single 8.5 draft was 8.0/9.0/9.0/8.0. The 0.8 separating the 4.0 drafts from the 4.8 drafts is source compliance alone, worth three tiers at a 25 percent weight, while information gain, specificity, and argument depth all stayed at their floor across those twenty drafts, according to the ForgeRank methodology page.

Which belief about AI content quality does the data actually support?

The common belief holds that AI drafts score poorly because the writing sounds synthetic. The evidence supports a different cause: they score poorly because the sourcing collapses, and the prose is a downstream symptom. The ForgeRank published dataset shows twenty drafts pinned at 4.0 and 4.8 with information gain, specificity, and argument depth all at their floor — a profile consistent with fluent writing and no verifiable claims underneath it. If synthetic phrasing drove the scores, those three dimensions would move together with source compliance. They did not move at all. Google Search Central's spam policies name scaled content abuse — mass-producing pages to manipulate rankings — as the target, regardless of whether AI or humans wrote them, which puts the enforcement weight on volume and value rather than authorship.

What if it doesn't work?

You cannot tell whether a source supports a claim. The source is paywalled or the link is dead. Treat an unreadable source as no source. Either find a replacement you can open, or rewrite the sentence as mechanism and remove the citation.

The draft is too long to check in one pass. Cap your claim list at the number you can verify in the time you have, and mark the rest as unverified in the document itself. The ForgeRank methodology page records that 13 of the 29 drafts hit the analysis cap, so an incomplete pass is the normal case, not a failure.

Every claim checks out but the draft still reads thin. Source compliance is one of four dimensions. A draft can pass sourcing and sit at the floor on information gain, which means the problem is that the sections restate what other pages already say.

A detector flags your own writing. Liang et al., arXiv 2304.02819, documented a roughly 61 percent false-positive rate on TOEFL essays from non-native English writers. A detector score is not evidence about authorship, and it is not evidence about sourcing either.

How long does this take?

Budget 30 to 60 minutes per 1,500 words for a first pass, and 15 minutes for a second pass on just the claims you flagged. The first pass is slow because you are opening sources. The second pass is fast because you already know which twelve sentences matter.

Finish checklist

  • Every number in the draft has a named source attached inline
  • Every named source resolves to something a stranger could find in two minutes
  • Every source you kept has been opened and read at the passage level
  • No sentence claims more than its source measured
  • Your claim list and the draft's claim count match

Frequently asked questions

Does a named organisation count as a source without a link?

Yes. A reader who can locate the organisation has what they need to verify. Add the link when you have it, but a missing URL does not make a correct attribution wrong.

How many claims in a typical AI draft need a source?

In the ForgeRank published dataset, 51 of 2,839 claims across 29 drafts were flagged as needing one, and 38 of those 51 had none attached.

Can I check sources without reading the full study?

No. Reading the abstract tells you what the authors wanted to claim. Reading the results section tells you what they measured, and the gap between those two is where misattribution lives.

What do I do with a statistic I cannot source?

Cut the number and keep the mechanism. "Inboxes overflow and the unread badge becomes a graveyard" survives without a figure. "The average person receives 120 emails a day" does not.

Is AI-written content penalised by Google?

Google Search Central's spam policies target scaled content abuse — mass-producing pages to manipulate rankings — regardless of whether AI or humans wrote them. Authorship is not the trigger; volume without value is.

You can now take any AI draft and separate the claims that hold from the claims that dissolve under a single question. The next step is small: open your most recent draft, number the paragraphs, and list every sentence containing a number or a named organisation. That list is your work queue, and it will tell you in ten minutes whether the draft is publishable or needs a rewrite.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.