AI content QA is a pre-publication review that scores a finished draft against fixed quality dimensions before anyone publishes it. It checks whether claims carry sources, whether the structure answers the query, and whether the draft clears a pass line. It does not detect authorship, and it does not fix prose sentence by sentence.

Why this matters

Three jobs get collapsed into one word. Authorship detection asks who wrote the text. Copy editing fixes commas and rhythm. Quality assurance asks whether the draft is true, sourced, and complete enough to publish. A team that buys a detector to solve a QA problem ends up with a green light and a page full of unsourced claims.

The ForgeRank methodology page describes a working version: each dimension is scored as the median of 3 samples, and the headline score is the average of the four dimensions. That structure exists because a single pass over a draft produces noise, and noise produces confident wrong verdicts.

The published evidence for why this matters sits in the ForgeRank published dataset (https://forgerankai.com/static/data/ai-content-quality-29-drafts.json). Across 29 drafts the scorer recorded 2,839 claims. Of those, 51 were judged to need a source and 38 had none attached. Three-quarters of the sourcing failures were caught by a reviewer reading the draft, not by a tool reading the prose.

What are the three jobs people confuse with AI content QA?

Authorship detection, copy editing, and quality assurance answer three different questions, and only one of them is about whether the draft is fit to publish. Detection asks who typed it. Editing asks whether it reads well. QA asks whether every claim survives a reader who checks.

Detection is the weakest of the three. Weixin Liang et al., in arXiv 2304.02819 (published in Patterns), found that roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated. Detector false positives fall disproportionately on non-native English writers. A gate that fails six in ten human essays from one population is not a gate.

Copy editing operates at the sentence. QA operates at the draft. A sentence can be clean and still assert a number nobody can verify. Teams that run editing without QA ship polished drafts with unfalsifiable claims, and the polish makes the claims harder to spot.

Where does AI content QA sit in the pipeline?

QA runs after drafting and before publication, over the whole draft rather than sentence by sentence. It is the last gate a draft passes, and it is the only gate whose output is a publish/hold decision rather than a list of edits.

Placement matters because of what QA needs to see. A claim-level check requires the full draft in front of the reviewer, so the reviewer can test whether the same figure appears twice with two different values, or whether a section's opening sentence answers the heading above it. Sentence-by-sentence review destroys that view.

The ForgeRank methodology page scores dimensions in tiers, with 4, 7, 8 and 9 all observed, which is why the headline score moves in steps. A tiered scale is a deliberate choice: it forces the reviewer to commit to a band instead of splitting hairs between a 6.8 and a 6.9. The cost is resolution. A draft that improves from 6.9 to 7.1 crosses the pass line, and a draft that improves from 5.0 to 6.9 does not.

What are the two gate lines, and what does each trigger?

A working QA process has two thresholds: a hard floor and a pass line. The floor catches drafts with structural or factual defects that no amount of editing fixes. The pass line separates drafts that ship from drafts that go back for rework.

The ForgeRank published dataset shows both lines in the distribution: 4.0 x10, 4.8 x10, 7.0 x3, 7.2 x4, 7.8 x1, 8.5 x1. Twenty of 29 drafts landed at or below 4.8, and none landed between 5.0 and 6.9. That empty band is the finding. Drafts either failed structurally or cleared the line; almost nothing sat in the middle waiting for a light edit.

The floor triggers a rewrite of the draft's claim layer: pull every number, attach a source or delete the number, then re-score. The pass line triggers a publish decision. A draft at 7.0 ships with the reviewer's notes attached for the next revision cycle. A draft at 4.8 does not ship at any level of copy editing, because the defect is in what the draft asserts, not how it sounds.

What do the four QA dimensions actually measure?

Each dimension measures one property of the draft, and each one stops applying in a specific case. Here is what the ForgeRank methodology page scores, and where each rule's authority ends.

Sourcing. It measures whether claims that require attribution have a named source attached. It matters because a claim without a source cannot be checked by the reader, and a reader who cannot check a claim has to decide whether to trust the writer. The consequence for the writer is that unsourced numbers get deleted, not softened. The boundary: a claim about the writer's own process needs no external source, because the writer is the source.

Accuracy. It measures whether the draft's stated facts match the sources it cites. It matters because a misquoted figure survives every downstream edit, since editors check grammar and reviewers check structure. The consequence is that the writer must open the source and read the sentence, not the abstract. The boundary: when the source itself is a summary of another source, cite the original or drop the claim.

Structure. It measures whether each section answers the question its heading poses, in the first paragraph under that heading. It matters because a reader who has to read three paragraphs to find the answer leaves. The consequence is that the writer writes the answer first and the context after. The boundary: narrative openings, where the answer is the story, are exempt.

Completeness. It measures whether the draft resolves the query it targets, including the follow-up questions the query implies. It matters because a draft that leaves the reader searching again has failed the reader and the page. The consequence is that the writer adds the two or three obvious follow-ups before publishing. The boundary: a draft scoped to one narrow question is complete when that question is answered, and padding it with adjacent topics lowers the score.

What do people believe about AI content QA that the data contradicts?

The common belief is that AI content QA is detection with a different label, and that clearing a detector means the draft is safe to publish. The evidence supports the opposite: detection is unreliable, and QA is about claims.

Two numbers settle it. Roughly 61 percent of human-written TOEFL essays were flagged as AI-generated, per Weixin Liang et al., arXiv 2304.02819. A gate with that error rate cannot be the thing standing between a draft and publication. Meanwhile the ForgeRank published dataset shows 38 claims that needed a source and had none attached, out of 2,839 total claims across 29 drafts. Those 38 are the actual risk surface, and no detector touches them.

The belief persists because detection produces a number, and a number feels like a gate. QA produces a decision, and a decision requires a reviewer to defend it. Teams reach for the detector because it is cheap and fast, then discover that a 98 percent human score tells them nothing about whether the pricing figure in paragraph four is real.

Google's position is on the record. Google Search Central's spam policies name scaled content abuse, mass-producing many pages to manipulate rankings, and the policy applies regardless of whether AI or humans wrote them. The March 2024 core update aimed to reduce low-quality unoriginal content in search results by 40 percent, per Google Search Central. Authorship is not the variable Google penalizes. Value is.

Key takeaways

  • AI content QA scores a finished draft against fixed dimensions and returns a publish/hold decision; authorship detection returns a probability and fails roughly 61 percent of human TOEFL essays, per Weixin Liang et al., arXiv 2304.02819.
  • The ForgeRank published dataset recorded 2,839 claims across 29 drafts, with 51 judged to need a source and 38 carrying none.
  • Two gate lines do the work: a floor that triggers a claim-layer rewrite, and a pass line that triggers publication. Twenty of 29 drafts in the ForgeRank published dataset sat at or below 4.8, and none landed between 5.0 and 6.9.
  • Google Search Central's spam policies target scaled content abuse regardless of whether AI or humans wrote the pages, and the March 2024 core update aimed to reduce low-quality unoriginal content in search results by 40 percent.

Frequently asked questions

Is AI content QA the same as AI detection?

No. Detection estimates who wrote a passage. QA scores whether the draft's claims are sourced, accurate, structured, and complete. A draft can score 100 percent human on a detector and still fail QA on 38 unsourced claims, which is what the ForgeRank published dataset recorded across 29 drafts.

How many samples does one QA score need?

The ForgeRank methodology page scores each dimension as the median of 3 samples, then averages the four dimensions into the headline score. Three samples per dimension exists to absorb reviewer variance. One pass over a draft produces a number that moves when the same reviewer reads it twice.

Why does the headline score move in steps?

The ForgeRank methodology page judges dimensions in tiers, with 4, 7, 8 and 9 all observed. Tiered scoring forces a band commitment instead of a hairline distinction between 6.8 and 6.9. The tradeoff is resolution: a 5.0 to 6.9 improvement does not cross the pass line, while 6.9 to 7.1 does.

Does AI content QA apply to human-written drafts?

Yes. The ForgeRank published dataset's 38 unsourced claims came from drafts, and the sourcing defect is independent of who typed the sentences. Google Search Central's spam policies state that scaled content abuse applies regardless of whether AI or humans wrote the pages, which puts the same QA burden on both.

What triggers a rewrite rather than an edit?

A score at or below 4.8 triggers a rewrite of the draft's claim layer, because the defect sits in what the draft asserts. Twenty of 29 drafts in the ForgeRank published dataset landed there. Copy editing cannot repair a claim with no source; only deleting the claim or attaching the source can.

The gate you actually need

AI content QA is a claims review with a threshold, run on a finished draft, before publication. It is not detection and it is not editing, and the teams that conflate the three ship drafts that pass a detector and fail a reader. The ForgeRank published dataset's empty 5.0-to-6.9 band is the clearest signal available: drafts fail structurally or they clear, and the middle is where wishful editing lives.

Tonight, open your last published draft and highlight every number in it. Any highlight without a source name attached gets deleted or sourced before you write another word.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.