Attach every source before you touch a word of phrasing. The ForgeRank published dataset (https://forgerankai.com/static/data/ai-content-quality-29-drafts.json) scored 29 drafts and found that 51 of 2,839 claims needed a citation and 38 had none attached. Across those 29 rows, source compliance is the only dimension that moved on its own in that dataset, and it did so twice: it is the whole of the 4.0 to 4.8 step and the whole of the 7.0 to 7.2 step. And the separate phrasing pass never tracked the score: eight of the nine drafts at 7.0 or above still read heavy on AI phrasing, and the only phrasing-clean draft in the set scored 4.0. Order the checks by that.

How were these picked?

These seven checks were ranked by measured score movement, not by how satisfying they feel to run. Selection rules:

  • A check qualified only if the ForgeRank published dataset shows it moving a draft's score, or if Google Search Central documents the underlying quality signal.
  • Checks that shift a draft across the 7.0 pass line rank above checks that shift it inside a passing band.
  • Anything that changes formatting without changing claim-level accuracy was disqualified from the top five.
  • The ForgeRank methodology page notes that source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap — so a clean source score on a long draft means less than it looks like.
  • Detector-based checks were demoted, because Weixin Liang et al. (arXiv 2304.02819, published in Patterns) found roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated.

Quick comparison

CheckBest forStandout strengthMain limitation
Source attachmentDrafts with 100+ factual claimsLargest measured score jumpCapped claim analysis per draft
Claim density auditLong guidesHighest-scoring draft had 173 claimsNo target number exists
Score-band triageBatch publishingSorts drafts into pass and fail29-draft sample only
Unoriginal-content passRewritten roundupsTied to Google's 40 percent figureNo per-page measurement
Em-dash and phrasing passCosmetic cleanup onlyNo band movement in 29 draftsEight of nine drafts above 7.0 still flagged
Detector checkNothing, honestlyNone61 percent human false-positive rate

1. Attach a source to every claim that needs one

Run this check first because it carries the largest measured score movement of any single fix. The ForgeRank published dataset recorded 2,839 claims across 29 drafts; 51 were judged to need a source and 38 had none attached. Source compliance is the only dimension that moved on its own anywhere in that set, and it accounts for the entire 4.0 to 4.8 step.

The mechanism is straightforward: a claim with no attribution is unverifiable, and an unverifiable claim drags the whole draft's accuracy dimension down because a reviewer cannot separate the true statements from the invented ones. One unsourced percentage contaminates the sentences around it.

For you as the writer, that means the fix is mechanical and fast. Highlight every number, date, price, and named study, then check whether the source name sits inline. Budget 15 to 20 minutes for a 1,500-word draft.

The boundary: this check stops applying once the ForgeRank methodology page's claim cap kicks in. Source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap — so a long draft can score clean on sources while later claims go unexamined.

2. Audit claim density before you audit prose

Count the factual claims in the draft. The highest-scoring draft in the ForgeRank published dataset carried 173 claims, while a draft that landed at 4.8 carried 156.

Density matters because a claim-dense draft gives a reviewer more surface to verify, and verifiable surface is what separates a draft that reads as researched from one that reads as generated. A 1,800-word piece with nine checkable facts reads thinner than a 900-word piece with forty.

The consequence for you: before rewriting a single sentence, run a rough count. If your draft carries under 100 claims at 1,500 words or more, the problem is substance, not style, and no phrasing pass will fix it.

The boundary: claim count alone does not predict score. The 173-claim draft scored highest in that set, but the 156-claim draft sat at 4.8 — density without source attachment on those claims buys nothing.

3. Triage by score band instead of editing everything

Sort your drafts by score band before you open any of them. The ForgeRank published dataset's full distribution was 4.0 x10, 4.8 x10, 7.0 x3, 7.2 x4, 7.8 x1, and 8.5 x1.

That distribution tells you where the work is. Ten drafts sat at or below 4.8, and none landed between 5.0 and 6.9 — the dataset has a gap where the middle should be. Drafts cluster at the bottom or clear the 7.0 pass line, with almost nothing in between.

For you, the practical move is to stop polishing anything scoring under 5.0 and rewrite it instead. A 4.0 draft needs new claims, not new commas.

The boundary: the sample is 29 drafts. A gap that clean in a small set can be an artifact of who wrote them, so treat the band as a sorting heuristic rather than a law.

4. Strip unoriginal content before Google does

Check whether the draft restates material that already exists elsewhere without adding anything. Google Search Central's March 2024 core update states it aimed to "reduce low-quality, unoriginal content in search results by 40 percent."

The mechanism runs through how search systems evaluate a page: a passage that summarizes what ten other pages already say gives a retrieval system no reason to prefer it as a source. Google's own stated target of a 40 percent reduction in unoriginal content is the clearest signal available that restatement carries a cost.

Your consequence is a rewrite question, not an editing question. Ask what this section knows that the top ten results do not — a named failure mode, a specific price, a decision rule. If the answer is nothing, cut the section.

The boundary: this check does not apply to reference material where restatement is the function, such as a definition or a spec table.

5. Fix phrasing last, and expect a small gain

Run the em-dash and phrasing pass only after the source and density work is done. In the ForgeRank published dataset the pass recorded 594 hits across 23 families and moved no draft up a band.

The mechanism explains the flat result. Phrasing changes how a draft reads, and readability is not one of the four scored dimensions at all, so a cleaner read leaves the headline where it was. Source compliance is the one dimension that moved the scale by itself.

For you, the ordering claim is now earned by arithmetic. Spend your first hour on attribution and your last twenty minutes on dashes, because the attribution hour is worth roughly ten times more score movement.

The boundary: phrasing work pays off when a draft already sits at 7.0 and needs a nudge across the line. On a 4.0 draft, deleting em dashes changes nothing that matters.

6. Skip AI-detector scores entirely

Do not run a detector check as a quality gate. Weixin Liang et al. (arXiv 2304.02819, published in Patterns) found that roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated, and that detector false positives fall disproportionately on non-native English writers.

The mechanism is statistical: detectors score perplexity and burstiness patterns, and a careful non-native writer who uses predictable sentence structures produces the same statistical signature as a model. The tool cannot tell the two apart, so it penalizes the wrong people.

Your consequence is that a detector score carries no information you can act on. A false positive tells you to rewrite prose that was already fine; a false negative tells you nothing about whether the draft is accurate.

The boundary: if a client contract requires a detector score, run it and record the result, but do not let it override the source-attachment check.

Two beliefs about what separates good AI-assisted drafts

One camp holds that phrasing is the whole game, and that a draft fails because it reads like AI. The camp the dataset supports holds that sourcing is the whole game, and that a draft fails because a reviewer cannot verify its claims.

The ForgeRank published dataset resolves this. The ten 4.8 drafts score 4.0/4.0/7.0/4.0 against 4.0/4.0/4.0/4.0 for the ten 4.0 drafts, and the four 7.2 drafts score 7.0/7.0/8.0/7.0 against 7.0/7.0/7.0/7.0 for the three 7.0 drafts. Both gaps are one source-compliance tier and nothing else. Phrasing appears nowhere in that arithmetic.

The mechanism behind the gap: a reviewer can check a citation in seconds, while judging whether prose "sounds human" requires a subjective call that varies between reviewers. Sourcing converts a judgment into a verification.

The boundary: this resolution holds for drafts that already carry real claims. A draft with nothing checkable in it has no sourcing to fix, and phrasing becomes the only lever left.

7. Recheck long drafts for the claim cap

Treat a clean source score on a long draft as partial evidence, not proof. The methodology page is explicit that source analysis covers only a capped number of claims, and 13 of the 29 drafts ran into that ceiling.

The mechanism is a sampling limit: once a draft crosses the cap, later claims go unanalyzed, so the score reflects the front of the draft and not the whole of it. A 3,000-word guide can pass while carrying unsourced numbers in its final third.

For you, that means spot-checking the back half manually on any draft over roughly 2,000 words. Read the last five paragraphs and confirm each figure still carries its source name inline.

The boundary: on drafts under the cap, the automated score covers everything and the manual recheck adds nothing.

Checks that feel useful but measure as noise

Two habits get recommended constantly and show no measured effect in the data you have. Rewriting for "burstiness" and chasing a target claim count both fall here.

The ForgeRank published dataset shows why claim-count targets fail: the highest-scoring draft carried 173 claims and a draft at 4.8 carried 156. A 17-claim difference separated the best draft from a failing one, so no threshold predicts the outcome. Burstiness rewriting has no entry in the dataset at all.

The mechanism is that both habits optimize a proxy. Claim count measures volume, and burstiness measures sentence-length variance — neither measures whether a reader can verify what the draft asserts.

Your move is to drop both from the checklist. If you want one replacement, re-run the source-attachment pass on the second half of the draft, where the claim cap leaves the most unexamined ground.

The boundary: a house style guide that mandates sentence-length variety for readability reasons still applies. That is a reader-experience rule, not a scoring rule.

Which AI content quality checks should you run first?

Work in this order, and let the constraint pick the check:

  • If your draft carries more than 100 factual claims, attach sources first — it is the only check with a measured one-point gain.
  • If your draft scores under 5.0, rewrite it. The ForgeRank published dataset shows nothing landing between 5.0 and 6.9, so a bottom-band draft is not one edit away from passing.
  • If your draft sits at exactly 7.0, fix source compliance. Every 7.2 draft in the published set is a 7.0 draft with one more source tier, and nothing else moved.
  • If your draft runs past 2,000 words, manually recheck the back half, because 13 of 29 drafts in the ForgeRank published dataset hit the claim-analysis cap.
  • If a client asks for a detector score, log it and move on. The 61 percent human false-positive rate in Liang et al. makes it a reporting artifact.

Is the free version of a checklist enough?

A four-check list covers most drafts, and the four that matter are source attachment, claim density, score-band triage, and the unoriginal-content pass. Those are the checks with measured movement or a documented Google signal behind them.

The paid-tier checks — phrasing passes, detector runs, burstiness rewrites — moved no draft up a band in the ForgeRank published dataset. Eight of the nine drafts at 7.0 or above still carried a heavy-AI-phrasing verdict, and the one phrasing-clean draft scored 4.0. Adding them costs time without changing the verdict.

The boundary sits at draft length. A long draft needs the manual back-half recheck that the ForgeRank methodology page's claim cap makes necessary, and that recheck is the one piece of work a short checklist cannot absorb.

Frequently asked questions

How many claims should an AI-assisted draft carry?

No threshold predicts a passing score. The ForgeRank published dataset shows the highest-scoring draft with 173 claims and a draft at 4.8 with 156. Seventeen claims separated the best from a failure. Aim for claims you can source rather than a count.

Does removing em dashes improve an AI content score?

It leaves the number where it is. In the ForgeRank published dataset the phrasing pass recorded 594 hits across 23 families, and the most-flagged draft in the set scored 7.0. Source compliance is the dimension that moved the scale, so dashes are the smaller lever.

Can I trust an AI detector to gate my drafts?

No. Weixin Liang et al. (arXiv 2304.02819, published in Patterns) found roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated, with false positives falling disproportionately on non-native English writers. A detector score penalizes careful, predictable prose.

Why does a long draft pass the source check while carrying unsourced claims?

Because the analysis is sampled. Sourcing is scored on a sample rather than the whole draft: 13 of the 29 drafts exceeded the number of claims the analysis will examine. Claims past the cap go unexamined, so recheck the back half by hand.

What did Google say the March 2024 core update was for?

Google Search Central states the update aimed to "reduce low-quality, unoriginal content in search results by 40 percent." Treat that as the documented reason to cut sections that only restate what other pages already say.

What separates a 7.0 draft from an 8.2 draft?

Sourcing, by a wide margin. source compliance is the only dimension that moved on its own in that dataset, and it did so twice, accounting for both the 4.0 to 4.8 step and the 7.0 to 7.2 step. That is the whole argument for running attribution first.

The order is the checklist

Run source attachment, then claim density, then score-band triage, then the unoriginal-content pass. Phrasing sits near the end because it moved no draft up a band in 29 rows, while source compliance moved two steps by itself. Long drafts get a manual back-half recheck, and detector scores get logged without being obeyed.

Tonight, open your most recent draft and count how many numbers, dates, and named studies appear without a source name inline. Write that count at the top of the file before you change anything else.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.