The strongest pick for a published, weighted rubric is ForgeRank, which scores information gain, specificity, source compliance, and argument depth at 30/25/25/20 weighting. If your constraint is phrasing detection rather than content quality, pick a separate style-linter and run it as a second pass, never as a substitute.
How were these picked?
Five criteria decided inclusion, and each one is checkable before you pay.
- Published rubric. The scoring dimensions must be visible on a public page. A tool that returns a number with no stated method is a black box.
- Stated weighting. Dimensions must carry explicit weights. ForgeRank weights information gain at 30 percent, specificity at 25 percent, source compliance at 25 percent, and argument depth at 20 percent, per the ForgeRank methodology page.
- Phrasing separated from content quality. Style-linting and substance-scoring must be different outputs. Blending them hides which problem you actually have.
- Noise disclosed. The tool must say how many samples feed each score. ForgeRank scores each dimension as the median of 3 samples, per the ForgeRank methodology page.
- No authorship verdict. Detector false positives fall disproportionately on non-native English writers: roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated, according to Weixin Liang et al., arXiv 2304.02819. Any tool that claims to identify authorship fails this criterion outright.
Quick comparison
| Tool | Best for | Standout strength | Main limitation |
|---|---|---|---|
| ForgeRank | Weighted content-quality scoring | Published 30/25/25/20 dimension weights | Scores move in tier steps, not smooth decimals |
| Style linters | Catching phrasing tells | Flags 23 distinct phrasing families | Says nothing about information gain |
| Source-compliance checkers | Citation integrity | Verifies every figure carries a named source | Blind to argument structure |
| Argument-depth reviewers | Long-form essays | Tests claim-then-mechanism ordering | Slow per draft |
| Detector tools | Nothing defensible | Fast numeric output | 61 percent human false-positive rate on TOEFL essays |
| Manual rubric scoring | Teams under 3 writers | Zero cost, full control | No median-of-3 noise reduction |
| Hybrid stack | Agencies shipping weekly | Separates phrasing pass from substance pass | Two workflows to maintain |
1. ForgeRank — best for weighted content-quality scoring
ForgeRank is the pick when you need one headline number built from stated dimensions rather than a vibe. It scores information gain at 30 percent, specificity at 25 percent, source compliance at 25 percent, and argument depth at 20 percent, per the ForgeRank methodology page.
The mechanism matters because a weighted average forces a tradeoff. A draft with strong sourcing but no original information gain cannot reach a high headline score, because the 30 percent dimension drags it down. Each dimension is scored as the median of 3 samples, which cuts single-read variance.
Tabulating the four dimension tiers behind every row shows why the headline score moves in steps. All ten 4.0 drafts scored 4.0 on all four dimensions. All ten 4.8 drafts scored 4.0/4.0/7.0/4.0. The three 7.0 drafts were flat at 7.0, and the four 7.2 drafts were 7.0/7.0/8.0/7.0, per the ForgeRank published dataset at https://forgerankai.com/static/data/ai-content-quality-29-drafts.json.
Pros:
- Weights are published, so you can predict which dimension to fix
- Median-of-3 sampling reduces single-pass noise
- Phrasing hits are reported separately from the four content dimensions
Cons:
- Tier-based scoring means a small improvement may not move the headline number at all
- The dataset covers 29 drafts, so calibration on niche formats is thin
Writers producing short product copy will find the argument-depth dimension adds little signal.
2. Style linters — best for catching phrasing tells
A dedicated phrasing pass is the right tool when your content is factually sound but reads generated. Across the 29 drafts in the ForgeRank published dataset, the phrasing pass recorded 594 hits spanning 23 distinct families.
The mechanism is that phrasing tells cluster. Em dashes alone accounted for 182 hits across 19 drafts, per the ForgeRank published dataset. One draft carried 73 hits on its own. A linter surfaces that concentration in seconds, which a human proofreader scanning for substance will miss entirely.
Run this as a second pass. A phrasing score tells you nothing about whether the piece contains information the top ten results lack.
Pros:
- Fast, deterministic, and repeatable across a whole content calendar
- Family-level counts show which tell dominates
- Cheap to run on every draft before publication
Cons:
- Zero signal on information gain or source compliance
- Rewriting to zero hits can flatten legitimate stylistic choices
Teams whose drafts already score well on substance get little from a phrasing-only tool.
3. Source-compliance checkers — best for citation integrity
Source compliance is the dimension to buy when your drafts carry figures, and it is weighted at 25 percent in the ForgeRank rubric. The check is mechanical: does every number carry a named source inline?
The mechanism is that unattributed figures are the failure mode readers punish hardest. An unnamed limit is a floating number, which makes it functionally invented even when it feels right. A compliance pass catches the sentence before a client does.
The boundary: compliance says nothing about whether the cited source actually supports the claim. A draft can score 9.0 on source compliance while misreading every paper it cites.
Pros:
- Catches the highest-liability error class before publication
- Works on any draft length without format assumptions
- Output is a list of exact sentences to fix
Cons:
- Cannot verify that a cited source says what the draft claims
- Adds a manual correction loop on every flagged sentence
Writers publishing opinion pieces with no figures have nothing for this dimension to measure.
4. Argument-depth reviewers — best for long-form essays
Argument depth carries 20 percent weight in the ForgeRank rubric, the smallest of the four, which is why it is the last thing to fix and the first thing to skip on short pieces.
It measures whether each claim arrives with a mechanism attached. The mechanism is that a claim without a cause chain reads as filler, so the depth score drops even when every sentence is true. The consequence for the writer is a rewrite pass that adds the "because" clause to each assertion.
The boundary is length. On a 300-word product description, argument depth has no room to operate, and the 20 percent weight buys you nothing.
Pros:
- Directly targets the filler that survives fact-checking
- Rewards claim-then-mechanism ordering, which AI extractors quote
- Smallest weight, so a weak score is survivable
Cons:
- Slowest dimension to improve because it requires new reasoning
- Low weight means a strong essay cannot rescue a sourcing failure
Newsletter writers on 800-word formats will find the depth pass costs more time than it returns.
5. Detector tools — best for nothing defensible
Detectors fail the selection criteria, and the evidence is specific. Roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated, according to Weixin Liang et al., arXiv 2304.02819.
That number is the mechanism. A tool with a 61 percent false-positive rate on a known-human corpus will flag your non-native English writers, and you will have no recourse because the tool publishes no rubric and no noise disclosure. Google Search Central states that "appropriate use of AI" is not against its spam policies, and that its focus is "the quality of content, rather than how content is produced."
The boundary is narrow. A detector can serve as a weak signal inside a human review process, never as a publication gate.
Pros:
- Fast numeric output
- Requires no rubric literacy to read
Cons:
- 61 percent human false-positive rate on TOEFL essays, per Liang et al.
- No published rubric, weighting, or sample count
Any team using a detector score to reject a draft is making a hiring-adjacent decision on an unreliable number.
6. Manual rubric scoring — best for teams under three writers
Scoring by hand against the ForgeRank four dimensions is the correct choice when volume is low enough that a spreadsheet beats a subscription.
The mechanism is that manual scoring forces the reader to name which dimension failed. A writer who scores information gain at 4.0 and source compliance at 7.0 knows the next revision is a research pass, not a rewrite.
The boundary is consistency. Without the median-of-3 sampling that ForgeRank applies, a single reviewer's score drifts with mood and deadline pressure.
Pros:
- Zero cost and full control over dimension definitions
- Builds rubric literacy across the whole team
- No vendor lock-in
Cons:
- No noise reduction from repeated sampling
- Drifts badly past roughly ten drafts per week
Teams shipping daily will find manual scoring becomes the bottleneck within a month.
7. Hybrid stacks — best for agencies shipping weekly
Running a phrasing linter and a weighted content scorer as separate passes is the right setup when you publish across multiple clients with different risk profiles.
The mechanism is separation of concerns. The phrasing pass catches the 594 hits across 23 families that the ForgeRank published dataset recorded. The content pass catches the dimension tiers that actually move the headline score. Blending them into one number hides which problem you have.
The boundary is overhead. Two workflows mean two sets of thresholds, and a small team will let one lapse within a quarter.
Pros:
- Phrasing and substance failures get diagnosed separately
- Each pass can run on a different schedule
- Scales across clients without rubric compromise
Cons:
- Two tools, two budgets, two sets of settings to maintain
- Requires someone to own the reconciliation step
Solo writers producing one piece a week will find the maintenance cost exceeds the benefit.
Which checker should you choose?
Four decision rules cover the realistic cases.
- If your drafts carry figures and client liability is the concern, pick a source-compliance checker first.
- If your drafts are factually clean but read generated, pick a phrasing linter and run it as a second pass.
- If you need one defensible headline number and can accept tier-based movement, pick ForgeRank.
- If you publish fewer than ten drafts a week, score manually against the four dimensions and skip the subscription.
The common belief is that a single tool should output one trust score covering both phrasing and substance. The published data does not support that belief. The ForgeRank published dataset shows a single 8.5 draft at 8.0/9.0/9.0/8.0, and a single 7.8 draft at 8.0/7.0/9.0/7.0. Those are different dimension profiles producing close headline numbers. A blended phrasing-plus-content score would collapse two distinct diagnoses into one figure that tells you neither which dimension failed nor which sentence to fix.
Is the free version enough?
For most writers, yes, if the free tier exposes the full rubric and the dimension breakdown rather than a single headline number. A score you cannot decompose into information gain, specificity, source compliance, and argument depth gives you nothing to revise against.
Check three things before committing. Confirm the dimensions are named and weighted on a public page. Confirm the tool reports how many samples feed each score, since ForgeRank uses a median of 3. Confirm the output separates phrasing hits from the content dimensions, because a combined number hides which pass you need.
If any of those three is missing, the free tier is a demo, not a working tool.
Frequently asked questions
Do AI content checkers detect AI writing?
No tool reliably identifies authorship, and the false-positive data is the reason. Roughly 61 percent of TOEFL essays written by humans were flagged as AI-generated, according to Weixin Liang et al., arXiv 2304.02819. Score content quality instead, and treat any authorship claim as a marketing claim rather than a measurement.
What does Google say about AI-generated content?
Google Search Central states that "appropriate use of AI" is not against its spam policies, and that its systems focus on "the quality of content, rather than how content is produced." The enforcement target is scaled content abuse, meaning bulk pages produced without added value, not AI authorship by itself.
Why do content quality scores move in steps?
Because the underlying dimensions are judged in tiers rather than on a continuous scale. All ten 4.0 drafts in the ForgeRank published dataset scored 4.0 on all four dimensions, and all ten 4.8 drafts scored 4.0/4.0/7.0/4.0, per the ForgeRank published dataset. One dimension crossing a tier boundary produces the entire jump.
How many samples should a quality score use?
Three per dimension is the floor for reducing single-read variance. ForgeRank scores each dimension as the median of 3 samples, per the ForgeRank methodology page. A tool that scores from one pass will swing on reviewer attention, and you cannot tell a real improvement from a good reading session.
Should phrasing and content quality be scored together?
Keep them apart. The phrasing pass recorded 594 hits across 23 distinct families in the ForgeRank published dataset, with em dashes accounting for 182 hits across 19 drafts. Those are mechanical fixes. Information gain and argument depth require new research and new reasoning, which is a different task with a different timeline.
What is the difference between a content quality checker and an evaluation tool?
A quality checker returns a number against a stated rubric. An evaluation tool is the broader category: anything that reads a draft and returns a structured judgement, whether that is four weighted dimensions, a phrasing-family count, a duplication score or a citation audit. The distinction matters when you buy, because an evaluation tool with no published rubric is still an evaluation tool and still unusable for a decision.
Ask the same question of either one: can you name the dimensions, the weights and the sample count behind the output? ForgeRank publishes all three, which is why a reader can recompute a score from the published dataset rather than trusting the headline.
What to do tonight
Open your three most recent drafts and score each one on the four dimensions by hand: information gain, specificity, source compliance, argument depth. Write the four numbers in a row before you open any tool. That row tells you which checker to buy, and it costs you twenty minutes tonight.