The order matters more than the checklist. In the ForgeRank published dataset (https://forgerankai.com/static/data/ai-content-quality-29-drafts.json), 20 of 29 AI-assisted drafts scored at or below 4.8 before publishing, and the gap between the 4.0 cluster and the 4.8 cluster came from source compliance alone. Check claims first, then gain, then specificity, then argument.

What you need before you start

You need the draft in a form you can mark up, not a form you can only read. Open it in a document you can comment on, or paste it into a plain text file and work there.

Have these ready:

  • The draft, at whatever length it currently is. Do not edit prose while you check claims; the two jobs fight each other.
  • A browser with a second tab open for source lookups.
  • A way to count: a spreadsheet column, a numbered list, or a counter. You will be counting claims twice.
  • Roughly 40 minutes for a 1,500-word draft. Sourcing takes the longest block of that time.
  • Prior knowledge of your own subject area. You cannot judge whether a claim needs a source if you do not know which claims are contestable.

By the end you will have a claim ledger, a list of unsourced assertions, and a decision on whether the draft goes out, gets rewritten, or gets cut. The procedure assumes you wrote or commissioned the draft and can change it.

Step 1: Number every claim in the draft

Read the draft once and tag each sentence that asserts something a reader could dispute. A claim is any sentence stating a fact, figure, causal relationship, or comparison. Tag them in order: C1, C2, C3, and so on.

This step exists because the later steps operate on a list, and you cannot build a list while reading for flow. In the ForgeRank dataset, the scorer recorded 2,839 claims across 29 drafts; 51 were judged to need a source, and 38 of those had none attached. That ratio is the whole reason this step comes first.

Work sentence by sentence. Mark narrative, opinion, and instruction as non-claims and move past them. When a sentence contains two assertions, split it into two numbered claims. When you finish, the draft should read as a numbered spine with prose around it.

You should end with a count you can write at the top of the file. If your count for a 1,500-word draft lands under 20, you are reading for comprehension instead of for assertion, and you need to go back through.

Step 2: Decide which claims actually need a source

Go through your numbered list and mark each claim as sourced, needs-source, or self-evident. The test is simple: a claim needs a source when a reader could reasonably ask "says who?"

This is where drafts separate. The ForgeRank methodology page (https://forgerankai.com/blog/ai-content-quality-data) notes that source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap — meaning the drafts had more contestable claims than the checker would score. A draft with 40 flagged claims and no sources is not a draft with a small problem.

Two worked rewrites from typical draft sentences:

Before: "Sourcing matters more than most writers think."

After: "A claim needs a source when a reader could ask 'says who?'"

Before: "Many teams struggle with this step."

After: "Thirteen of the 29 drafts in the ForgeRank dataset hit the claim cap before the checker finished."

The before versions fail because they assert a quantity with no instance behind it. The after versions name the draft, the number, or the test.

Mark your list. You should be able to read the needs-source column and see the draft's actual risk profile.

Step 3: Attach a source to every claim you flagged

For each claim in the needs-source column, either attach a named source inline or rewrite the sentence so it no longer asserts the unsourced thing.

The mechanism is straightforward: an unsourced claim transfers risk from the writer to the reader, and readers resolve that risk by leaving. Google Search Central's March 2024 core update stated its aim as reducing low-quality, unoriginal content in search results by 40 percent, and unsourced assertion is one of the patterns that update targeted.

Attach sources by naming the publisher in the sentence: "Google Search Central's March 2024 core update..." or "the ForgeRank published dataset..." Keep the source name attached to the figure. A number without its source is a number you cannot defend.

Where you cannot find a source, rewrite. "Most teams run two tools" becomes "a two-tool stack shows up when one department adopts first and nobody consolidates." The rewrite keeps the insight and drops the unverifiable quantity.

When you finish, every claim on your list should be sourced, rewritten, or cut. The needs-source column should be empty.

Step 4: Score information gain on the whole draft

Ask one question of the draft as a unit: what does this contain that the top results on the topic do not?

The mechanism behind this check is that a page summarizing other pages adds nothing a retrieval system can prefer it for. Google Search Central's spam policies name scaled content abuse as generating content primarily to manipulate rankings without adding value — and a draft that restates the consensus is the same failure at smaller scale.

The consequence for you is a hard decision. If the draft's only claim to value is coverage, it needs a section that reports something first-hand: a number you measured, a failure mode you observed, a decision rule tied to a specific situation.

The boundary: this check does not apply to reference pages, glossaries, or internal documentation, where completeness is the value. It applies to anything published to earn search traffic.

Step 5: Check specificity sentence by sentence

Read the draft again looking only at nouns. Every claim sentence should carry at least one of: a number, a named entity, a version, a date, or a quoted line.

Specificity is what makes a claim checkable, because a reader can verify "38 of 51 flagged claims had no source" and cannot verify "sourcing is important." In the ForgeRank dataset, the drafts that cleared the pass line pulled ahead on more than one dimension at once — the single 7.8 draft scored 8.0, 7.0, 9.0, and 7.0 across the four dimensions, and the single 8.5 draft scored 8.0, 9.0, 9.0, and 8.0. The gains came from several dimensions moving, not one.

Two consequences follow from that. First, a draft with strong sourcing and weak specificity stalls around the same score as a draft with the reverse — you cannot trade one for the other. Second, the fastest lift on a specific draft comes from whichever dimension is lowest, not from the one you enjoy editing.

The boundary: specificity has a ceiling. A draft stuffed with numbers that do not bear on the argument reads as noise, and the numbers stop doing work.

Step 6: Test the argument's depth

Read the draft's main claims and ask whether each one explains why, not just what.

The mechanism: a claim with a mechanism attached survives a skeptical reader because the reader can follow the causal chain and check each link. A claim without one collapses the moment a reader disagrees, because there is nothing underneath it to examine.

The consequence is that you rewrite flat assertions into causal ones. "Attach sources" becomes "attach sources because an unsourced figure is a figure the reader cannot verify, and unverifiable figures are what the March 2024 core update was built to demote."

The boundary: not every sentence needs a mechanism. Instructions, definitions, and transitions carry the reader forward and do not need to justify themselves.

Step 7: Run the false-positive check on your own draft

Before you trust any automated score, remember what automated detectors get wrong. Weixin Liang et al., in arXiv 2304.02819 (published in Patterns), found that detectors flagged roughly 61 percent of TOEFL essays written by non-native English speakers as AI-generated. That figure is the reason a detector score is evidence, not a verdict.

The common belief is that a low AI-detection score means the draft is safe to publish. The evidence in the ForgeRank dataset points the other way: 20 of 29 drafts scored at or below 4.8 on quality dimensions, and none landed between 5.0 and 6.9 — the drafts that failed, failed on sourcing, gain, specificity, and argument, not on detector output. A draft can pass every detector and still carry 38 unsourced claims.

The resolution: use detectors to flag passages for a second read, and use the claim ledger to decide whether the draft ships. The two tools answer different questions.

What if it doesn't work?

You cannot find a source for a claim you believe is true. Rewrite the sentence to state the mechanism instead of the quantity. "Most drafts fail on sourcing" becomes "a draft fails when a flagged claim carries no named source." The rewrite is weaker as a claim and stronger as a fact.

Your claim count is far higher than the checker's cap. In the ForgeRank dataset, 13 of 29 drafts hit the cap. When you hit it, stop counting and start cutting. A draft with 60 flagged claims is a draft that needs a shorter scope, not a longer source list.

Every claim is sourced and the draft still reads flat. Sourcing is one dimension. Check information gain and specificity next — the 7.8 and 8.5 drafts in the dataset moved on multiple dimensions at once, and a single-dimension fix caps out.

A detector flags a passage you wrote yourself. Run the Liang et al. figure against your instinct: roughly 61 percent of human-written TOEFL essays drew false positives. Read the passage for the actual problem — vague qualifier, unsourced quantity, missing mechanism — and fix that.

You cannot tell whether a sentence is a claim. Ask whether a reader could respond "prove it." If yes, it is a claim.

How long does this take?

Budget 40 minutes for a 1,500-word draft, split roughly as 15 minutes on Steps 1 through 3, 15 minutes on Steps 4 through 6, and 10 minutes on Step 7 and the checklist. Sourcing is the longest single block, and it is the block that moves the score most.

Finish checklist

  • Every claim in the draft carries a number, a named entity, or a source.
  • The needs-source column is empty.
  • At least one section reports something the top results do not contain.
  • No sentence asserts a quantity without a named source.
  • The draft has been read once for claims and once for prose, in that order.

Frequently asked questions

Can I check AI content with a detector alone?

No. A detector returns a probability, not a quality score. In the ForgeRank dataset, 20 of 29 drafts scored at or below 4.8 while the failures traced to sourcing, gain, specificity, and argument — dimensions a detector does not measure.

What is the single highest-leverage check?

Source compliance. In the ForgeRank dataset, the 0.8 gap between the 4.0 and 4.8 clusters came from source compliance alone at a 25 percent weight, while information gain, specificity, and argument depth sat at their floor across those 20 drafts.

How many claims should a 1,500-word draft have?

More than you expect. The ForgeRank dataset recorded 2,839 claims across 29 drafts, and 13 of those drafts exceeded the checker's per-draft claim cap. Count first, then decide what to cut.

Does Google penalize AI-written content by itself?

No. Google Search Central's spam policies target scaled content abuse — producing content primarily to manipulate rankings without adding value. AI authorship is not the trigger; unoriginal, valueless output is.

Should I fix the lowest-scoring dimension first?

Yes. The dataset's top drafts moved on several dimensions at once, and the 8.5 draft scored 8.0, 9.0, 9.0, and 8.0 — no single dimension carried it. Fix the floor before you polish the ceiling.

Open your current draft and number the first ten claims in it. If two of them have no source attached, you have found the work for tonight.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.