We scored 29 AI-assisted drafts before publishing them, on four weighted dimensions: real information gain (30%), specificity density (25%), source compliance (25%), and argument depth (20%). The headline score is the average of those four, rounded to one decimal. Same rubric, same three-sample median per dimension, same week in September 2026.

What this sample is: all 29 drafts are long-form article drafts for one platform (Google/Blog). Median length 11,867 characters, range 2,755 to 22,352. None of them are short-form posts or video scripts, so nothing here should be read as a claim about how social formats score. They also skew toward SEO, marketing and technology topics, because that is what our corpus holds.

Where the numbers come from: every figure in the distribution, sourcing and coverage tables below was recomputed from our own stored scoring runs, and the per-draft dataset is published: one row per draft, with platform, length, headline score, the four dimension tiers, phrasing families and claim counts. The tier analysis in the next section is arithmetic on those same rows, so you can reproduce it. Every figure on this page comes from that file.

We are publishing this because the shape of the distribution contradicts the editing advice we kept hearing. The common line is that AI drafts are "a bit weak" and need a polish pass. Our scores looked nothing like that, and once we broke the scale down by dimension, the reason turned out to be mechanical.

The scores form a staircase

Here is the full distribution of the 29 drafts:

Score bandDraftsWhat the notes said
4.0 (all four dimensions at the floor)10Consensus restatement: advice that is correct, widely repeated, and adds nothing a reader could not get from the first search result.
4.8 (three dimensions at the floor, source compliance higher)10Same emptiness, but no fabricated claims: nothing invented, nothing sourced either.
7.0–7.2 (competent)7At least one real mechanism with a "why", some checkable specifics, still padded.
7.8–8.5 (distinctive)2Non-consensus mechanism, named sources, concrete numbers, argument that builds.

In our published 29-draft dataset, twenty of twenty-nine drafts landed below 5.0 and none landed between 5.0 and 6.9. A draft that only lacks polish should sit in the middle of the scale. The middle stayed empty.

Every step up that staircase came from one dimension

This is the part we did not expect. The four dimensions are scored in tiers, so the headline number moves in steps. When we tabulated the tiers behind each step, each step turned out to be a single dimension changing. Nothing else moved.

HeadlineDrafts Information gain
weight 30%
Specificity
weight 25%
Source compliance
weight 25%
Argument depth
weight 20%
What moved
4.0104.04.04.04.0baseline
4.8104.04.07.04.0Source compliance only (+3 tiers × 25% = +0.75)
7.037.07.07.07.0baseline
7.247.07.08.07.0Source compliance only (+1 tier × 25% = +0.25)
7.818.07.09.07.0Information gain and source compliance
8.518.09.09.08.0Specificity and argument depth

Read the two bolded rows again, because they are the whole finding. All ten drafts at 4.0 have the same four tiers: 4.0, 4.0, 4.0, 4.0. All ten drafts at 4.8 have 4.0, 4.0, 7.0, 4.0. The 0.8 that separates our worst drafts from our second-worst tier is entirely one dimension: whether the sentences that asserted something about the world could be traced to anything.

Nothing about the writing improved. In those twenty drafts, information gain never moved, specificity never moved, and argument depth never moved. The only difference was that the 4.8 group avoided claims it could not support, while the 4.0 group made them anyway. Twenty of twenty-nine drafts sat at or below 4.8, and across those twenty the source-compliance tier was the only tier that ever left the floor.

The same pattern repeats one step higher, which is why we treat it as a pattern. Between 7.0 and 7.2, among drafts that already had a real mechanism and checkable specifics, the only tier that moved was source compliance again, from 7.0 to 8.0. Competent-plus-sourced scored higher than competent, with nothing else changed.

Only above that do the content dimensions start doing work. The single 7.8 draft got there on information gain (7.0 to 8.0), with source compliance already at 9.0. The single 8.5 draft, the only one that cleared our own 8.5 publish bar, moved on specificity (7.0 to 9.0) and argument depth (7.0 to 8.0), while information gain stayed at 8.0 and source compliance stayed at 9.0. Both drafts at 7.8 or above had source compliance at 9.0.

So the practical order we would give anyone is the reverse of the usual advice: sourcing is the gate, and substance is the ceiling. A draft that cannot attribute its claims is stuck near the floor no matter how well it is written. A draft that can attribute them has cleared the gate and is then limited only by how much it adds.

Phrasing is measured separately from quality

Alongside the four dimensions, each draft runs through a phrase-pattern pass that counts how many distinct AI-tone families it contains: em-dash overuse, hedge density, buzzwords and padding connectors, "not X — it's Y" constructions, and so on. Across the 29 drafts this pass recorded 594 hits spanning 23 distinct families.

Phrase familyTotal hitsDrafts containing itMedian hits per draft that had any
Em-dash overuse182197
Hedge-word density ("often", "arguably")83252
Buzzwords & padding connectors ("unlock", "game-changer")49232
Negation contrast ("not X — it's Y")49143.5
"Not X. Y." ellipsis47213
Semicolon-heavy sentences45142.5
…17 further families, 139 hits between them

The most-flagged draft in the set carried 73 hits. It scored 7.0. The draft with the single largest em-dash count, 58 of them, also scored 7.0. Meanwhile the two best drafts in the set carried 28 and 25 hits, which is squarely mid-range, and eight of the nine drafts scoring 7.0 or above still received a "heavy AI phrasing" verdict.

The cleanest draft in the entire set, the only one that came back with a clean verdict, scored 4.0. It was tidy, and it said nothing.

Taken together, a low phrasing count tells you very little about the score. That is why we treat tone as a separate cosmetic pass and report it separately instead of folding it into the quality number. It is also why "remove the AI tells" makes a poor quality strategy. You can strip every em dash from a draft and still have a draft that restates the first search result.

Claim count predicts quality no better

We counted the claims in each draft and compared the count to the score. The counts ranged from 25 to 173 across the set, and the range was just as wide at the bottom as at the top:

Score bandDraftsClaims detected (range)
4.01043–149
4.81025–156
7.0–7.2747–164
7.8–8.52123–173

The highest-scoring draft in the set also carried the most claims (173), and a draft that landed at 4.8 carried 156. The shortest claim list in the whole set (25) also scored 4.8. Claim volume tells you how long and how dense the piece is. What separated the top drafts was whether a reader could check any of it, and that has nothing to do with how much they said.

What "needs a source" looked like

The source-compliance tier comes from a measurement. The draft is annotated sentence by sentence, and each sentence that asserts something gets three ordered questions: is this claim checkable against the outside world, does it require a source, and is a source attached? Across the 29 drafts, that pass read 2,839 claims. 51 required a source, and 38 of those 51 had nothing attached at all.

The pattern is easier to see in one piece. One 22,352-character draft in the same set carried 164 claims, 9 of which required a source and 6 of which had none. The flagged sentences are statistics, specific figures, and named-entity claims: the sentences a reader would try to verify first, and the ones that damage trust fastest when they cannot be checked.

In practice you either attach evidence or you stop making the claim. On one long draft from this set we worked through the flagged sentences and resolved each one of them: attaching a traceable source wherever one existed, and rewriting the sentence so it no longer asserted an unsourced fact wherever none did. That is the whole repair procedure, and it is mechanical.

The machinery behind these numbers

Since this page asks you to trust measurements, here is what produced them, at the level of mechanism.

  • Sentence-level claim annotation. The source tiers above come from asking each asserting sentence three ordered questions: is it checkable against the outside world, does it require a source, and is one attached. Two of those combinations are contradictory, because a sentence that cannot be checked against the outside world has no external source to require. Those get corrected instead of counted, so the counters never hold impossible labels.
  • Phrasing is counted per family. That is why the table above has 23 rows rather than one number, and why a draft can be both "heavy on AI phrasing" and high-scoring. Phrasing and content quality are deliberately kept apart.
  • A claim cap that limits the annotation only. Long drafts are capped at 100 annotated claims per piece. That cap bounds how much of a long draft gets sentence-level sourcing analysis. It never touches the content score, which always reads the full text. In this set 13 of 29 drafts hit the cap, so per-draft sourcing coverage ranged from 57.8% to 93.5% (both figures are per row in the dataset). We would rather state the cap and the coverage than report a clean 100%.
  • Readiness checks sit outside the score. Format length limits, fabricated-source patterns, and arithmetic consistency inside the text are checked independently, because a draft can score well on the four dimensions and still be unpublishable on a specific platform.

You can point any of this at a draft you already wrote, including one you wrote somewhere else entirely, and get the same four dimensions, the same separate phrasing read, and the same weakest-dimension call. That is the read the numbers on this page came from.

Where this analysis is limited

Five limits worth stating plainly, because a data post that hides its limits is just marketing with numbers:

  • One platform, one format. All 29 drafts are long-form article drafts for a single platform. Nothing on this page is evidence about short-form posts, threads, carousels or video scripts, where both the rubric shape and the failure modes are different.
  • Sample bias. These are drafts from our own corpus, not a random sample of the internet. They skew toward SEO, marketing and technology topics.
  • Coverage caps. On long drafts, sourcing analysis runs on a capped number of claims: 100 per piece. In this set 13 of 29 drafts hit that cap, so coverage ranged from 57.8% to 93.5%. The tail of a very long piece is not fully analyzed, and we say so on the report.
  • Five drafts carried no pass line. In this set 24 of 29 rows carry a publish bar of 8.5, and 5 rows carry none, so for those we can describe the score but not their distance from a bar. Their industry label is also unset. Both appear as empty fields in the dataset.
  • These numbers are first-party. Every figure here is our own measurement of our own drafts, published so it can be checked row by row. There is no independent replication. Treat this as a published lab notebook, rather than as external evidence.

What we do not claim

No ranking promises, no traffic predictions, no authorship detection. A score describes a draft. It says nothing about how a search engine or an audience will respond. Two drafts with the same score can perform differently for reasons that have nothing to do with the text. A high score also says nothing about whether the topic was worth writing about. That choice stays with you.

How we measured

Each draft was read three times by the same rubric and the median was kept, per dimension, to reduce single-run noise. The four dimensions are scored independently and in tiers rather than on a continuous scale, which is why the headline number moves in steps such as 4.0 to 4.8 to 7.0, and why the staircase table above can attribute each step to a single dimension. The headline score is the average of the four dimension scores, rounded to one decimal. Runs were stored with their timestamp and configuration, and the figures on this page were recomputed from those stored records rather than retyped. The per-draft dataset is published here: 29 drafts, one row each.

If you want the same read on your own draft, you can run one piece free, no signup, at forgerankai.com. You get the score, the weakest dimension, the top issue, and the phrasing patterns. Whether you agree with the verdict is the useful part.

How to cite this data

The dataset is published with a version number and a citation string, so it can be referenced instead of paraphrased. The current release is v1.0.0, and it carries the field definitions, the scoring method and the known limitations alongside the rows. It contains the derived measurements only, with no draft text, so citing it does not redistribute anyone's writing.

Cite as: ForgeRank AI. (2026). AI Content Quality: 29 AI-Assisted Drafts Scored Before Publishing, v1.0.0. forgerankai.com/static/data/ai-content-quality-29-drafts.json

Two limits matter when you reuse these figures. Scores are not reproducible to a single point, because the same text scored twice can differ by 0.3 to 0.5, so a published score is a point on a range rather than a fixed value. And the set contains only AI-assisted drafts, with no human control group, so it cannot support any claim about how often detectors flag human writing.

The four checks you can run without any tool

  1. Read your first sentence. If it could open any article on this topic, it carries no information. Name a number, a name, a date, or a decision.
  2. Give every number an owner. Write down what each figure belongs to: which study, which sample, which date. If you cannot fill that line, delete the number. This is the dimension that separated our 4.0 drafts from our 4.8 drafts, and it is the one you can fix without rewriting anything else.
  3. Separate phrasing from substance. Clean tone last. Decide what the piece adds, then remove the filler around it. Expect no band movement from the tone pass, because in this data it never produced any on its own.
  4. Count what you added. Highlight anything a reader could not have gotten from the first search result. If nothing is highlighted, the draft is a summary, not a piece.
Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.