Paste your draft into a checker that scores four separate dimensions, read the dimension that moved, then spend 20 minutes fixing the claims, phrasing, and sourcing the tool cannot judge. A checker that returns one number tells you nothing actionable. Four numbers tell you which specific habit is dragging the draft below the 7.0 pass line.

Who this is for and what you'll have at the end

This is for writers, editors, and content leads who already publish regularly and want a repeatable quality gate before a draft ships. It assumes you can open a JSON file in a browser or spreadsheet and read a score table; no coding is required. You will finish with a scored draft, a list of flagged claims, and a manual pass that catches what no automated checker can score. Budget 20 minutes for the manual work on top of whatever the tool takes. The reference point throughout is the ForgeRank published dataset (https://forgerankai.com/static/data/ai-content-quality-29-drafts.json), which scores 29 drafts across four dimensions.

Before you start: what you need

  • A finished draft, at least one full pass of your own editing behind it.
  • Access to a checker that reports four separate dimension scores rather than a single aggregate.
  • The ForgeRank published dataset open in a second tab so you can compare your draft against real scored examples.
  • The ForgeRank methodology page open, because it documents the claim cap that shapes what the tool can and cannot see.
  • A text editor with find-and-replace, for the phrasing pass.
  • 20 minutes of uninterrupted time for the manual review.

Step 1: Score the draft on all four dimensions

Run the draft through the checker and record four numbers, not one. The four dimensions are the same ones the ForgeRank published dataset tabulates: a topic score, a structure score, a phrasing score, and a sourcing score. Every row in that dataset is a four-number profile, which is why the tiers separate so cleanly.

The mechanism matters because a single aggregate hides the shape of the problem. In the dataset, every 4.0 draft scored 4.0 on all four dimensions, and every 4.8 draft scored 4.0/4.0/7.0/4.0. The 4.8 drafts were not uniformly better; one dimension had moved. When you see four numbers, you know which one to work on.

The consequence for you: a draft sitting at 6/10 with a 9 on structure and a 3 on sourcing needs sourcing work, not rewriting. Chasing the aggregate score wastes the session.

The boundary: this stops applying when you are scoring a draft that is not prose, such as a spec sheet or a data table. The four dimensions assume sentences with claims in them.

Step 2: Read the distribution, not your gut

Compare your four-number profile against the dataset's tiers before you change a word. The dataset's distribution is narrow and instructive: ten drafts flat at 4.0, ten at 4.0/4.0/7.0/4.0, three flat at 7.0, four at 7.0/7.0/8.0/7.0, one at 8.0/7.0/9.0/7.0, and one at 8.0/9.0/9.0/8.0.

The mechanism: each step up that distribution is a single dimension moving, because the drafts that scored well did not fix everything at once. The 7.8 draft lifted structure to 9.0 while leaving phrasing at 7.0. The 8.5 draft lifted both structure and phrasing and still left topic at 8.0.

For you, that means the fastest gain is the dimension with the most headroom, and the ceiling is set by whichever dimension you have not touched.

The boundary: the distribution describes this dataset's 29 drafts, so treat it as a reference shape for comparison. It does not predict what your specific draft will score after edits.

Step 3: Separate the writing checker from the detector

Run two tools, not one, because a writing quality checker and an AI detector measure different properties and disagree on the same file. The writing checker reads properties of the text: claim density, sourcing, phrasing patterns, structure. The detector reads authorship signals and returns a probability.

The disagreement is documented and severe. Weixin Liang et al., in arXiv 2304.02819 (published in Patterns), found that roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated. Those essays were written by people. The detector was reading a signal that correlates with non-native English phrasing, not with machine authorship.

Your move: when the two tools disagree, trust the writing checker for editorial decisions and treat the detector score as noise unless you have a specific reason to care. A detector flag on a human-written draft is a known failure mode, not a verdict.

The boundary: this stops applying if your publishing contract or client explicitly requires a detector threshold. Then the detector is a compliance gate, and you clear it as a separate task from improving the writing.

Step 4: Audit the flagged claims against their sources

Open the sourcing list and check every claim the tool flagged as needing a source. The dataset recorded 2,839 claims across the 29 drafts; 51 were judged to need a source and 38 of those had none attached. That is a 75 percent miss rate on the claims that mattered most.

The mechanism: an unsourced claim is the single easiest thing for a reader or an AI system to discount. The ForgeRank methodology page notes that source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap, so the flagged list is a floor, not a ceiling.

Fix each one by attaching the source inline, naming what the number belongs to. A price, a percentage, or a limit is only checkable if the reader knows what it applies to.

The boundary: drop the claim rather than dress it up if you cannot find a real source. An unattributed statistic is worse than no statistic.

Step 5: Run the phrasing pass and count the families

Search the draft for the phrasing patterns the tool flags, and count distinct families rather than total hits. Across the 29 drafts the phrasing pass recorded 594 hits spanning 23 distinct families. Em dashes alone accounted for 182 hits across 19 drafts, and the most-flagged single draft carried 73 hits.

The mechanism: 73 hits in one draft means the same habit repeated, which is what a reader feels as a machine rhythm even when no single sentence is wrong. Twenty-three families means the checker is looking for a wide set of patterns, and your draft will trip only some of them.

For you, the fix is targeted: find your top two families and cut them, then re-score.

The boundary: the phrasing pass measures pattern frequency, so a draft that deliberately uses one pattern for effect, such as a repeated structural device in a listicle, will register hits that are not defects.

Step 6: Do the 20-minute manual pass

Automation cannot score three things, and you cover them by hand. Read the draft aloud and mark every sentence where your voice drops out. Check that each section's first paragraph answers its own heading in 40 to 60 words, because that is the passage a retrieval system lifts. Confirm that every claim you kept in Step 4 now carries its source name inline.

The mechanism: a checker scores properties it can measure, and answer-first placement, voice consistency, and source-to-claim proximity are semantic judgments. Google Search Central's March 2024 core update states its aim was to reduce low-quality, unoriginal content in search results by 40 percent, and unoriginal summarising is exactly what a property-based checker cannot detect.

The boundary: if the draft is under 800 words, the answer-first pass takes five minutes and the rest of the manual review still applies.

Two beliefs about quality checkers, and what the data says

The common belief is that a checker's job is to predict whether AI systems will penalise the draft, so you optimise for the score. The belief the evidence supports is that a checker's job is to tell you which specific habit is weak, and the score is a byproduct.

Resolve it against the published dataset. If the score were the point, the ten drafts at 4.0 and the ten at 4.8 would be treated as different quality tiers. They are not: the 4.8 drafts are identical to the 4.0 drafts on three dimensions and differ on one. The dataset's structure argues for reading the profile, not the total.

What if it doesn't work?

The tool returns one score with no breakdown. You are using an aggregate checker. Switch to one that reports the four dimensions separately, or you cannot act on the result.

The detector flags a draft you wrote yourself. This is the documented failure mode from Liang et al. Check the writing checker's dimension scores instead, and only chase the detector if a contract requires it.

Every claim comes back flagged and you cannot fix them all. The methodology page notes source analysis is capped per draft, so 13 of the 29 drafts in the dataset hit that cap. Prioritise the claims carrying numbers; drop the ones you cannot source.

The phrasing pass flags 70-plus hits. You have one dominant family, not 23 problems. Find the top family and cut it, then re-run.

The four scores are flat at 4.0. You are at the dataset's floor. Pick the single dimension with the most headroom and fix only that, because the distribution shows drafts move one dimension at a time.

How long does this take?

The scored run takes a few minutes. The manual pass takes 20 minutes on a draft under 2,000 words. The claim audit is the variable: with 38 of 51 flagged claims missing sources in the dataset, expect the sourcing step to consume most of your session.

Finish checklist

  • Four dimension scores recorded, not one aggregate.
  • Your profile compared against the dataset's tiers.
  • Detector score logged separately and not used as an editorial verdict.
  • Every kept claim carries its source name inline.
  • Top two phrasing families cut and the draft re-scored.
  • Answer-first opening verified for each section.

Frequently asked questions

Does a writing checker replace an AI detector?

No, and they measure different things. The writing checker reads properties of the text: claims, sources, phrasing, structure. The detector reads authorship signals. Liang et al. found roughly 61 percent of human-written TOEFL essays were flagged as AI-generated, which is why the detector score should not drive editorial decisions.

Why did the checker flag a claim I know is true?

A true claim with no source attached still reads as unsourced. The dataset recorded 51 claims judged to need a source, and 38 had none attached. Attach the source inline and name what the number belongs to; the flag clears.

Can I fix a low score by adding length?

No. The dataset's tiers separate on dimension movement, not word count. The 8.5 draft scored 8.0/9.0/9.0/8.0, which is a profile, not a length. Adding words without moving a dimension leaves the score where it was.

What is the pass line?

The 7.0 tier is the reference floor in the dataset, with three drafts flat at 7.0 and four at 7.0/7.0/8.0/7.0. A draft scoring 6/10 sits below that line, and the fix is the dimension with the most headroom.

How many phrasing families should I expect?

The dataset recorded 594 phrasing hits spanning 23 distinct families. Your draft will trip a subset. Find the two families with the most hits, cut them, and re-score rather than chasing all 23.

You now have a scored draft, a sourced claim list, and a manual pass that covers the semantic judgments a checker cannot make. The next step is to re-run the four-dimension score and compare it against the tier you started at, so you can see which dimension actually moved. Open the ForgeRank published dataset side by side with your new profile tonight and match your numbers to the closest tier.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.