A claim-first review works like this: read for claims and their sources before you touch the prose, then structure, then phrasing last. The ForgeRank published dataset (https://forgerankai.com/static/data/ai-content-quality-29-drafts.json) scored 29 drafts across four dimensions, and every step up the score distribution came from a single dimension moving. Reviewing in that order catches what style-first reading buries.

Who this is for and what the job produces

You review AI drafts before they publish. You don't need a QA tool, a plagiarism checker, or a style guide. You need 40 minutes and a spreadsheet row per draft.

The output is a review record: one line per flagged sentence, each line carrying the sentence, the flag, and the fix. Hand it back and the writer knows exactly what to change. No rewrite, no vague notes.

Time: 35 to 45 minutes for a 1,500-word draft. Longer if the draft carries more than 150 claims.

What do you need before you start?

  • The draft in a format you can edit or annotate — Google Docs, Word, or a plain text file.
  • A four-column spreadsheet: sentence, dimension, flag, fix.
  • The source list the writer used, if one exists. If it doesn't, that's your first finding.
  • One hour of uninterrupted time. Claim checking breaks badly when interrupted.

Step 1: Count the claims before you read for quality

Open the draft and mark every sentence that asserts a fact, a number, a date, a price, or a causal claim. Label each one with a number in the margin.

This matters because claim count sets your review budget. The ForgeRank published dataset recorded 2,839 claims across 29 drafts, and the highest-scoring draft carried 173 claims while a draft that landed at 4.8 carried 156. A draft with 170 claims needs a different review than one with 40, and you can't know which you have until you count.

Count in one pass, no editing. If the draft runs past 150 claims, you're reviewing a research piece and should budget two sessions. The ForgeRank methodology page notes that source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap — so a long draft can outrun any automated check you point at it.

When you finish, you should have a numbered list in the margin and a total. If the total is under 30, the draft is thin on substance and you can flag that before you read a single sentence for quality.

Step 2: Check every number against its source before you read anything else

Take your numbered list from Step 1 and work down it. For each claim, ask one question: does the source actually support this sentence?

The ForgeRank published dataset found that of 2,839 claims across 29 drafts, 51 were judged to need a source and 38 had none attached. That's a 75 percent miss rate on the claims that mattered most — the ones a reviewer flagged as needing backing. You find those 38 by checking, not by reading.

Two failure modes live here. A claim with no source attached, and a claim whose source doesn't say what the sentence says. The second is worse. A writer who cites a real study for a conclusion the study never reached has produced something that reads authoritative and fails on contact.

Work claim by claim. When a number has no source, write "no source" in the flag column. When a source exists but doesn't support the claim, write "source doesn't support" and quote the sentence the source actually makes.

The pass worked when every claim in your numbered list has either a source name and a page reference, or a flag. Claims with neither are the ones that get published and cause corrections.

Step 3: Check the structure against the questions a reader would ask

Read only the headings, in order, without the body text. Write down the question each heading answers.

A draft that survives claim checking can still fail here. Headings that describe topics rather than answer questions push the reader to search again, which is the failure mode Google's helpful-content guidance names directly. Google Search Central states that its systems aim to reward content "written by people, for people" — and a heading like "Considerations" tells a reader nothing about whether their question gets answered.

Fix by rewriting each heading as the literal question. "Pricing" becomes "How much does it cost?" "Implementation" becomes "How long does setup take?"

The ForgeRank published dataset shows why this step sits third rather than first. Structure is cheap to fix — a heading rewrite takes 90 seconds — while an unsupported claim takes a source hunt or a deletion. Spend your attention on the expensive errors first; the cheap ones will still be there.

Step 4: Run the phrasing pass last, and count the hits

Search the draft for the AI phrasing families: em dashes, "not X but Y" pivots, ordinal openers, hedge words, and two-beat closing couplets. Tally each family.

The ForgeRank published dataset recorded 594 phrasing hits across 23 distinct families in those 29 drafts. Em dashes alone accounted for 182 hits across 19 drafts, and the most-flagged draft carried 73 hits. That last number is the one to sit with — 73 phrasing flags in a single draft means the writer has a habit, not a slip.

Run this pass with a search function, not by eye. Search the em dash character and count. Search "not just" and "isn't about" and count. Search "often," "typically," "generally" and count.

The boundary on this step: phrasing flags don't block publication the way unsupported claims do. A draft with 20 em dashes and every claim sourced is publishable after a light edit. A draft with clean prose and 38 unsourced claims is not.

Step 5: Fill in the review record and hand it back

Build one row per flagged sentence in your four-column spreadsheet: sentence, dimension, flag, fix. Send it as the review, not as notes attached to a rewrite.

The four dimensions in the ForgeRank dataset are claims, sources, structure, and phrasing, and the tier data shows they don't move together. All ten drafts that scored 4.0 scored 4.0 on all four dimensions. The ten drafts at 4.8 scored 4.0/4.0/7.0/4.0 — one dimension moved and the rest held. The three drafts at 7.0 were flat at 7.0. The four at 7.2 sat at 7.0/7.0/8.0/7.0, the single 7.8 draft at 8.0/7.0/9.0/7.0, and the single 8.5 draft at 8.0/9.0/9.0/8.0.

That distribution is the argument for the read order in this guide. A draft doesn't climb by getting better at everything. It climbs when one dimension moves and the rest hold, which means the reviewer's job is to find which dimension is holding the draft down and name it precisely enough that the writer can move it.

Your record is done when every row has a fix the writer can apply without asking you a question. "Fix the tone" fails that test. "Delete the em dash in sentence 14" passes it.

What if it doesn't work?

The draft has more than 170 claims and you can't finish in one sitting. Split by section and review each section's claims independently. The highest-scoring draft in the ForgeRank published dataset carried 173 claims, so a draft at that size is a full review, not a skim.

A claim has a source but you can't verify the source exists. Flag it as unverifiable and hold the draft. A source you can't find is a source the reader can't find either.

The writer pushes back on a "source doesn't support" flag. Quote the exact sentence the source makes and put it next to the draft sentence. The gap is usually visible in two lines.

You flagged 60 phrasing hits and the draft still reads flat. Phrasing flags don't measure readability. Read the first paragraph of each section aloud; if you run out of breath, the sentences are too long and that's a separate fix.

The draft scores well on claims and sources but readers bounce. Check the headings against Step 3. Answer-first headings change what a reader does in the first five seconds, and no amount of claim accuracy fixes a heading that doesn't promise an answer.

Should you review for style first?

No. Style-first reviewing buries the two errors that matter: a claim with no source, and a claim the source doesn't support. Both are invisible to a phrasing pass, and both are the errors that produce public corrections.

The ForgeRank published dataset makes the case numerically. Of 2,839 claims, 51 needed a source and 38 had none — and the phrasing pass found 594 hits across the same 29 drafts. Phrasing problems are loud and numerous. Source problems are quiet and few. A reviewer who starts with the loud problem spends their attention on the cheap fix and ships the expensive one.

Finish checklist

  • Every claim in the margin list has a source name and reference, or a flag.
  • Every heading reads as a question a reader would actually type.
  • Phrasing hits are tallied by family, with the em dash count written down.
  • The review record has one row per flagged sentence, each with an applicable fix.
  • No row says "improve" or "tighten" without naming the sentence and the change.

Frequently asked questions

How long should a review take?

35 to 45 minutes for a 1,500-word draft under 150 claims. Past 150 claims, budget two sessions — the ForgeRank published dataset's top draft carried 173 claims, and source checking scales linearly with claim count.

Can I use an automated checker instead?

Use it for the phrasing pass only. The ForgeRank methodology page notes that source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap — so automated source checking stops before your draft does.

What's the single most common failure?

A claim that needs a source and doesn't have one. The ForgeRank published dataset recorded 38 such claims out of 51 flagged, which is the highest-leverage thing a reviewer can find.

Do phrasing hits block publication?

No. They're an edit, not a gate. A draft with 20 em dashes and every claim sourced publishes after a light pass; a draft with clean prose and unsourced claims doesn't publish at all.

How do I know when the review is finished?

When every row in your record carries a fix the writer can apply without asking you a question. If any row needs a conversation, the review isn't done.

You can now run a claim-first review and hand back a record a writer can act on without a follow-up meeting. Tonight, take one AI draft you've already published and count its claims — just the count, nothing else. If it comes in under 30, the thin reading was never a mystery.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.