Every AI-generated number needs four checks before you publish it: what it measures, which population it belongs to, what year it describes, and how large the sample was. A draft can pass every aggregate quality score and still carry a wrong sentence. The ForgeRank published dataset (https://forgerankai.com/static/data/ai-content-quality-29-drafts.json) shows this gap at scale.

What you need before you start

This guide assumes you can open a spreadsheet, read a JSON file, and follow a citation to its origin. No statistics training is required. Budget 45 to 90 minutes for a 1,500-word draft, longer if you have to chase down primary sources.

Bring these:

  • The draft with every numeric claim still in place, including inline citations
  • A spreadsheet with four columns: claim text, what it measures, population, year, sample size
  • Access to the original source for every cited figure, not the aggregator that repeated it
  • The ForgeRank published dataset open in a second tab for comparison
  • A calculator or spreadsheet formula for the arithmetic self-check

By the end you will have a claim-by-claim audit sheet, a list of numbers you can defend, and a list you had to cut or rewrite. The four dimensions below are the same ones an expert reviewer applies when scoring a draft.

Step 1: Extract every numeric claim into a spreadsheet

Open the draft and copy each sentence containing a number, percentage, dollar figure, date, or sample size into column A of your spreadsheet. One claim per row, quoted exactly as written.

This step matters because aggregate scores hide sentence-level failures. The ForgeRank published dataset recorded 2,839 claims across 29 drafts; 51 of those were judged to need a source, and 38 had none attached. That is 38 sentences a reader could pull apart. An aggregate score never shows you which ones.

Work through the draft top to bottom. Do not skip a number because it looks obvious. "Roughly half of informational searches" is a claim. "Most teams" is a claim. If it quantifies anything, it goes in the sheet.

When you finish, count your rows. If your 1,500-word draft produced fewer than 40 claims, you have missed some. The highest-scoring draft in the ForgeRank published dataset carried 173 claims, and a draft that landed at 4.8 carried 156. Dense numeric writing is normal; if your sheet looks thin, re-read the draft with a highlighter.

A common failure here: people extract only the numbers with a percent sign. Dates, sample sizes, and dollar figures get left behind, and those are exactly the claims that break under scrutiny.

Step 2: Ask what each number actually measures

For each row, write in column B the specific quantity the number describes. "Engagement" is not a measurement. "Click-through rate on the first email in a sequence" is.

This matters because the mechanism that produces a number determines what you can say about it. A figure about open rates tells you nothing about revenue, because the two are measured at different points in the same funnel and move for different reasons. When you write a sentence that swaps one for the other, you have not exaggerated a statistic; you have invented a new one.

The consequence lands on you as the writer. If your sentence says a metric "drives" an outcome and the underlying study only measured correlation at one point in time, a reviewer who checks the source will find the mismatch. That is the failure that costs credibility, and it survives every aggregate score.

The boundary: this rule stops applying when the source itself states a causal design. A randomized controlled trial that assigned participants to conditions can support a causal sentence. An observational survey cannot, no matter how large the sample.

Step 3: Identify which population the number belongs to

Write in column C the group the number describes: which country, which industry, which job role, which age band, which platform. Then compare it to the group your sentence implies.

Population mismatch is the quietest error in AI-assisted drafts, because the number is real and the sentence reads cleanly. A figure collected from enterprise marketing teams does not describe solo consultants. A study of US college students does not describe working adults in Germany. The number survives; the sentence does not.

Your move as the writer is either to narrow the sentence to match the population or to drop the figure. Narrowing is cheaper. "Among enterprise teams surveyed" costs you four words and buys you a defensible claim.

The boundary case: if the source explicitly states that its sample was drawn to be nationally representative, you can generalize to that nation's adult population. Absent that statement, the number belongs to the sample and nothing wider.

Step 4: Pin down the year the data describes

Column D gets two dates: the year the data was collected and the year the source was published. They are frequently different, sometimes by three or four years.

The mechanism is simple. A statistic is a photograph of one moment. Pricing changes, platform rules change, and market composition changes. A 2021 figure about remote work adoption describes a world with different office norms than 2026. When you present it as current, you have made a claim the source never made.

For you, this means every time-sensitive number needs its collection year in the sentence or the citation. If you cannot find the collection year, treat the figure as undated and either cut it or label it as such.

The boundary: some measurements are stable enough that age matters less, such as a physical constant or a documented platform rule that has not changed. For anything tied to behavior, pricing, or adoption, the year is part of the claim.

Step 5: Check the sample size behind each figure

Column E holds the sample size and the sampling method. A percentage without an n is a rumor with a percent sign.

The mechanism runs through the arithmetic of uncertainty. A finding from 40 respondents carries a wide margin of error; the same percentage from 4,000 respondents carries a narrow one. When you quote the number without the sample, you are presenting the precision of the second case while holding the evidence of the first.

Your consequence: any figure where you cannot locate the sample size gets flagged in the sheet and either cut or rewritten without the number. Vague-but-true beats precise-but-invented.

The boundary: government censuses and full-population administrative datasets have no sampling error, because they measured everyone. That is the only case where you can skip the sample-size column.

Step 6: Run the arithmetic self-check

Sum your percentage columns and reconcile your stated totals against the underlying counts. If four categories are listed at 28%, 31%, 22%, and 14%, the total is 95%, and the missing 5% needs an explanation or the numbers are wrong.

This step catches a specific failure: fabricated or garbled figures that pass every qualitative read. The ForgeRank published dataset's own arithmetic is a useful model. Of 2,839 claims, 51 were judged to need a source and 38 had none attached, which means 13 of the 51 were properly cited. Those three numbers reconcile: 38 plus 13 equals 51.

Two more checks belong here. First, confirm that your stated total matches the count of items you listed; a sentence claiming "four dimensions" must name four. Second, check whether your source analysis hit a cap. The ForgeRank methodology page (https://forgerankai.com/blog/ai-content-quality-data) notes that source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap. If your own audit caps out, say so rather than implying full coverage.

If the percentages do not sum and you cannot find the missing share, the figure is not usable. Cut it.

Step 7: Resolve the two rival beliefs about AI detection

The common belief holds that an AI detector score settles whether a draft is trustworthy. The evidence supports a different conclusion: detector output and factual accuracy are separate measurements, and detector output carries its own error rate.

Weixin Liang et al., arXiv 2304.02819, published in Patterns, found that roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated. That is a false-positive rate landing disproportionately on non-native English writers. A detector flag is not evidence of a fabricated statistic, and a clean detector score is not evidence that your 38 unsourced claims are fine.

Resolve it against the ForgeRank published dataset. Twenty of twenty-nine AI-assisted drafts scored before publishing landed at or below 4.8, and none landed between 5.0 and 6.9. The full distribution was 4.0 ten times, 4.8 ten times, 7.0 three times, 7.2 four times, 7.8 once, and 8.5 once. Notice what that distribution means: the pass line sits at 7.0, and the cluster below it is where the unsourced claims live. The detector question and the sourcing question have different answers, and only one of them is about whether your numbers are true.

What if it doesn't work?

The source link is dead or paywalled. Move to the primary source the aggregator cited. If the trail ends, the figure is unverifiable and gets cut. Do not substitute a similar-looking number from a different study.

The number appears in ten articles but no original study. This is citation drift: each article copied the previous one. Trace back until you find the first source. If none exists, treat the figure as unsourced regardless of how many sites repeat it.

Your percentages sum past 100%. Overlapping categories are the usual cause, such as respondents who could select multiple answers. Rewrite the sentence to say so, or convert to a count instead of a share.

The sample size is buried in a methods appendix. Search the source for "n =", "participants", or "respondents". If the appendix is inaccessible, flag the claim and either drop the number or describe the finding without it.

Your audit sheet stops growing because you hit a tool's claim cap. Note the cap in your write-up. The ForgeRank methodology page documents this exact limit, and pretending you covered everything is worse than stating your coverage.

How long does this take?

A 1,500-word draft with 60 numeric claims takes 45 to 90 minutes for a full pass, faster once the spreadsheet columns are set up. The slow part is chasing primary sources, not filling cells. Budget extra time for any figure where the population or collection year is not stated in the source itself.

Finish checklist

  • Every numeric claim sits in its own spreadsheet row with a quoted sentence
  • Columns B through E are filled for each claim, or the claim is flagged
  • Percentages in each section sum to 100%, or the gap is explained
  • Every figure carries its source name and collection year inline
  • Unverifiable claims are cut, not softened

Frequently asked questions

Why can a draft pass every aggregate check and still be wrong?

Aggregate scores average across sentences, so 38 unsourced claims can hide inside 2,839 total claims. The ForgeRank published dataset shows exactly this: 51 claims were judged to need a source, and 38 had none attached. An aggregate number never tells you which sentences failed.

What does the four-dimension check actually test?

It tests what a number measures, which population it describes, what year it covers, and how large the sample was. A figure can be real and still fail all four when your sentence implies a different quantity, group, year, or precision than the source supports.

Is a clean AI detector score proof my statistics are sound?

No. According to Weixin Liang et al., arXiv 2304.02819, published in Patterns, roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated. Detector output measures writing style, and it carries a false-positive rate that falls hardest on non-native English writers.

How do I handle a statistic I can't trace to any original study?

Cut it. Citation drift means the number may have been copied across dozens of articles with no primary source behind it. Repetition is not verification, and a figure you cannot trace is a figure you cannot defend when someone asks where it came from.

What does Google say about AI-generated content and search ranking?

Google Search Central states: "Our focus is on the quality of content, rather than how content is produced." That means AI authorship alone is not the issue. Unsourced or inaccurate statistics are the quality problem, and they affect how the page performs.

You now have a repeatable audit: extract every claim, test it against the four dimensions, run the arithmetic, and resolve the detector question separately from the sourcing question. The next step is mechanical. Open your most recent draft tonight, paste every numeric sentence into a spreadsheet, and count how many rows have an empty population column. That count is your real quality score.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.