A 7.0 headline score is the ship line and 4.8 is the kill line. ForgeRank scored 29 AI drafts against four dimensions, each dimension the median of three samples, the headline the average of those four. Nothing landed between 5.0 and 6.9. Ten drafts sat at 4.0, ten at 4.8, nine at or above 7.0.

Why does scoring beat reading your draft again?

Reading your own draft measures fluency, and fluency is the one thing AI already delivers. The ForgeRank dataset shows the distribution is bimodal: 20 of 29 drafts scored at or below 4.8, and none landed between 5.0 and 6.9. A draft either clears the bar or misses it by a wide margin, which means a gut check tells you almost nothing.

The full distribution was 4.0 x10, 4.8 x10, 7.0 x3, 7.2 x4, 7.8 x1, 8.5 x1, according to the ForgeRank published dataset. That gap between 4.8 and 7.0 is where the decision lives. A 4.8 draft reads fine. So does a 7.2 draft. The difference sits in claims, sources, and structure, none of which your eye catches on a third pass at midnight.

Headline scoreDraftsBand
8.51ships
7.81ships
7.24ships
7.03at the ship line
4.810below
4.010below
5.0–6.90empty

The middle of the scale stayed empty. Twenty of the 29 drafts landed at or below 4.8 and none landed between 5.0 and 6.9, so a half-fixed draft does not drift toward the ship line: it stays in the low band until something a reader can verify gets added. The rows behind these counts are published at ai-content-quality-29-drafts.json.

The cost of skipping the score lands later. Google's March 2024 core update aimed to reduce low-quality, unoriginal content in search results by 40 percent, per Google Search Central. A 4.8 draft that ships is a bet against that number.

What do the four dimensions actually measure?

Each dimension is scored as the median of 3 samples, and the headline score is the average of the four, per the ForgeRank methodology page. Dimensions are judged in tiers (4, 7, 8 and 9 all observed), which is why the headline score moves in steps rather than sliding smoothly.

Claim density. This dimension measures how many checkable assertions a draft makes per section. The mechanism that makes it matter is that a draft with few claims gives a reader nothing to verify, because an unsourced paragraph that asserts nothing specific cannot be wrong and cannot be trusted either. For the writer, the consequence is that a thin draft scores low no matter how clean the prose reads. The rule stops applying when a section is deliberately transitional: a two-sentence bridge between arguments carries no claims and does not drag the score.

Source attachment. This dimension measures what share of claims that need a source actually carry one. The mechanism is that a reader who cannot open a link has to take your word, because there is no path from your sentence back to a primary document. For the writer, the consequence is arithmetic: across those 29 drafts the scorer recorded 2,839 claims, 51 were judged to need a source, and 38 had none attached, per the ForgeRank published dataset. The boundary is that not every claim needs a citation. A statement of mechanism ("the median of three samples reduces variance") is self-evidencing and needs no link.

Structural extractability. This dimension measures whether each section resolves one question on its own. The mechanism is that a section which depends on the paragraph above it cannot be lifted out and quoted, because the meaning breaks at the seam. For the writer, the consequence is that a well-argued point buried mid-article contributes nothing when a reader or a crawler pulls a single block. The rule stops applying to narrative openings, where a scene earns its place by setting up the problem the next section solves.

Voice regularity. This dimension measures repeated sentence shapes, stock transitions, and punctuation habits. The mechanism is that a pattern repeated across 2,000 words reads as generated, because human writers vary rhythm without deciding to. For the writer, the consequence is measurable: in the ForgeRank published dataset, phrasing moved no draft up a band, while source compliance moved the scale twice. Eight of the nine drafts scoring 7.0 or above still read heavy on AI phrasing. The boundary is that a single stylistic choice (one em dash, one short punch sentence) is invisible to the scorer. Only repetition registers.

Two beliefs about scoring, tested against the data

The common belief is that a high score means the draft is good writing. The rival belief, the one the dataset supports, is that a high score means the draft is checkable. The highest-scoring draft in that set carried 173 claims, and a draft that landed at 4.8 carried 156, per the ForgeRank published dataset. Nearly the same claim count, a 3.7-point gap.

The difference sits in what happened to those claims. Source analysis runs on a capped number of claims per draft, per the ForgeRank methodology page, so 13 of the 29 drafts hit that cap. A draft that hits the cap has more claims than the scorer will examine, which means the ceiling on its score is set by how many of the examined claims carry sources, not by how much it says. Writing more does not raise the score. Attaching a source to what you already wrote does.

How do you run the score yourself in twenty minutes?

Work in one pass per dimension and resist fixing prose until the numbers are in. Open the draft and count claims in the first three sections. Mark each one that a skeptical reader would want to verify. Then check which of those marked claims carry a link or a named source, and count the gap.

The 7.0 ship line is the number to hold. Below it, the draft goes back for sources, not for rewriting. Between 5.0 and 6.9 is empty territory in the dataset, which tells you a draft that feels "almost there" is actually sitting at 4.8 and needs the same fix as a draft that feels weak. Above 7.0, stop editing structure and check voice: read the piece aloud and listen for the third time you used the same sentence shape.

What most people get wrong about scoring AI content

The mistake is treating an AI-detection score as the gate. Detector false positives fall disproportionately on non-native English writers, per Weixin Liang et al., arXiv 2304.02819 (published in Patterns), and roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated in that study. A tool that flags human writing at that rate cannot decide whether your draft ships.

Google's spam policies name scaled content abuse, mass-producing many pages to manipulate rankings, regardless of whether AI or humans wrote them, per Google Search Central spam policies. Authorship is not the violation. Volume without value is. That is why the four-dimension score is the right gate: it measures the thing the policy actually targets, which is whether each page carries claims a reader can check.

Key takeaways

  • A 7.0 headline score is the ship line; 4.8 is the kill line, and nothing in the ForgeRank dataset landed between 5.0 and 6.9.
  • Each dimension is the median of three samples, so a single strong section cannot carry a weak draft.
  • Across 29 drafts, 51 claims needed a source and 38 had none attached, per the ForgeRank published dataset.
  • Phrasing moved no draft up a band in the published set; source compliance moved two of the four steps by itself.
  • Google's March 2024 core update targeted a 40 percent reduction in low-quality, unoriginal content, per Google Search Central.

Frequently asked questions

What score should AI content hit before publishing?

Hold 7.0 as the ship line. The ForgeRank dataset shows nine of 29 drafts sat at or above 7.0, and the jump from 4.8 to 7.0 is where the decision sits. A draft below 7.0 goes back for sourced claims, not for a prose rewrite.

Is AI detection the same as a quality score?

No. Detector false positives fall disproportionately on non-native English writers, per Weixin Liang et al., arXiv 2304.02819, where roughly 61 percent of human-written TOEFL essays were flagged. A quality score measures claims and sources; a detector measures surface patterns.

Why does the score move in steps instead of sliding?

Dimensions are judged in tiers (4, 7, 8 and 9 all observed), per the ForgeRank methodology page, so the headline score is the average of four tier values. That is why you see 4.0, 4.8, 7.0, and 7.2 rather than a smooth curve.

How many claims should a draft carry?

The highest-scoring draft in the ForgeRank set carried 173 claims. A draft that landed at 4.8 carried 156. The count is close, so the score turns on how many of those claims carry a source, not on writing more.

Does removing em dashes really change the score?

In the ForgeRank published dataset the phrasing pass recorded 594 hits across 23 families and still moved no draft up a band, while source compliance moved two steps by itself. Fix sources first; punctuation is the smaller lever.

Open your last AI draft and count the claims in the first three sections, then mark which ones carry a link. If the gap is wider than three, the draft is sitting at 4.8 and the fix is a source, not a rewrite.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.