Evaluation goes wrong in a predictable way. Someone runs a detector, gets a number, and treats that number as a verdict on the work. The number answers a narrower question than the one being asked, and four separate checks get collapsed into one. Separating them is most of the job.

Four lanes exist here, and each one closes a question the others leave open. Detection asks whether a model was involved. Quality asks whether the piece is worth publishing. Evidence asks whether the individual claims hold. Tone asks whether the prose reads as machine-made. Run them out of order and you spend effort where it changes nothing.

Lane one: detection answers a question about the tool, not the text

A detector scores how statistically typical a passage looks to a language model. GPTZero-style tools compute perplexity and burstiness, where burstiness is the variance in how surprising each sentence is. DetectGPT (Mitchell et al., arXiv 2301.11305, ICML 2023) tests a related property by perturbing a passage and measuring whether its likelihood moves. Machine-written text sits near a local optimum of the model's own probability, so perturbation lowers the score. Human text sits away from that optimum and stays roughly flat.

Both methods describe the shape of the passage in front of them. Neither has access to your revision history, your sources, or the fact that you typed the sentences. Liang et al. (arXiv 2304.02819, published in Patterns) found that roughly 61 percent of TOEFL essays written by non-native English speakers were flagged as AI-generated. Those writers used no model. They wrote in the controlled register an exam rewards, and regularity is the property being measured.

The boundary matters more than the reading. A high detector score says nothing about whether the content is weak, and a low score says nothing about whether it is strong. The lane closes one question and leaves the publish decision untouched.

Lane two: quality is the lane that maps to a publish decision

Quality evaluation reads four properties and weights them: information gain at 30 percent, specificity at 25 percent, source compliance at 25 percent, and argument depth at 20 percent. Information gain asks what the reader learns beyond what they arrived with. Specificity asks whether claims carry numbers, names, or instances. Source compliance asks whether the figures trace to something the writer opened. Argument depth asks whether the piece explains a mechanism rather than stacking parallel points.

Run against the ForgeRank published dataset, the tiers separate cleanly and show how the checks relate to one another. The ten drafts scoring 4.0 were flat at 4.0 on all four dimensions. The ten scoring 4.8 held 4.0 on information gain, 4.0 on specificity, and 4.0 on argument depth, with source compliance at 7.0. The four 7.2 drafts read 7.0/7.0/8.0/7.0, and the single 8.5 draft read 8.0/9.0/9.0/8.0. Across the whole set, source compliance was the only dimension that moved a score on its own, and it moved it twice.

That is the argument for reading quality as four numbers rather than one. A draft can improve on the dimension that decides publication while the other three sit still, and a blended total hides which one moved. It also hides the reverse case, where a draft holds its total while trading one dimension for another.

Specificity is the easiest of the four to see in practice once you look for instances rather than claims. A sentence saying that many teams struggle with review gets replaced by the instance: the four drafts at 7.2 that scored 7.0 on argument depth because each one listed three practices without ordering them or explaining which failure each practice prevents. The second version can be argued with and the first cannot.

Lane three: evidence is checked claim by claim, and it is where coverage runs out

Evidence work is enumeration and attribution. Every factual claim gets listed, and each one either traces to a source or it does not. The ForgeRank published dataset records 2,839 claims across those 29 drafts, averaging 97.9 claims per draft, with the longest draft carrying 173.

Enumeration is also where a tool has to admit a limit rather than report full coverage. Source analysis runs on a capped number of claims per draft, and 13 of the 29 drafts hit that cap. The cap limits annotation rather than scoring, which cuts both ways: the score still reads the whole draft, while the per-claim breakdown stops at the cap. A clean source score on a long draft therefore means less than it appears to, and the ForgeRank methodology page documents the cap rather than hiding it.

The lane closes the question of whether a specific number can be traced. It leaves open whether the source is any good, and whether the claim was worth making in the first place. A figure can be perfectly attributed and still be the wrong figure for the argument.

Lane four: tone is a separate signal, and editing it does not move quality

Tone work counts phrasing patterns: em-dash density, hedge density, buzzword families, padding connectors, and constructions like the reversal that announces a contrast it never delivers. In the 29-draft set this pass recorded 594 hits across 23 distinct families, with em-dash density alone firing 182 times across 19 drafts.

The tempting move is to treat tone as the whole job, clean the phrases, and call the draft fixed. The dataset closes that door. Cleaning AI-sounding phrasing moved no draft up a score band. Eight of the nine drafts scoring 7.0 or above stayed flagged as heavy on AI phrasing after cleanup. The most-flagged draft carried 73 hits and scored 7.0. The draft with 58 em dashes also scored 7.0. The one draft whose phrasing came back clean scored 4.0.

Read those together and the two lanes stop being interchangeable. Tone responds to editing while leaving whether the draft is worth publishing where it was. A writer who spends an afternoon trimming em dashes has changed how the prose reads, and the phrasing counter is the only measurement that moved.

The order that wastes the least effort

Evidence comes first. It is the cheapest of the four lanes and it is the only dimension in the published set that ever moved a score by itself, so attaching sources to numbers is where an hour goes furthest.

Quality comes next, because that reading is what a publish decision actually consults. Run it after the sources are in, since source compliance is one of the four inputs to the total.

Tone follows. It is real work and it is measurable, and its ceiling is cosmetic in the sense the numbers show: it can make a draft read differently without making it better.

Detection goes last, and only when the setting demands it, such as a submission that will be screened. Inside a publishing workflow it answers a question nobody was asking.

One page covers the boundary between the first lane and the third: detection is not fact checking sets out why a detector reading and an evidence check are different instruments. If you want the shorter version of the quality lane, how to score a draft before publishing walks the four rows. For the evidence lane specifically, checking sources and finding unsupported claims cover the two halves, and what a false positive means picks up the first lane in detail.

Applying the lanes to a draft you did not write

Client submissions and inherited drafts run through the same four lanes, with one adjustment: the evidence lane moves first and becomes the gate rather than a score. A draft whose numbers trace nowhere cannot be repaired by better prose, so the attribution question gets settled before anyone edits a sentence.

Quality comes second and answers a different question than the author expects. Authors ask whether the draft is good. The useful reading is which of the four dimensions is carrying it and which one is holding it down. A submission that scores 8.0 on three dimensions and 4.0 on source compliance has one problem, and it is addressable. The same 8.0 total built from four 8.0s has a ceiling that is harder to argue about.

Tone comes third, and for inherited work it is mostly a diagnostic rather than a task. A high phrasing count on a draft with strong evidence points at a writer who drafted fast and edited for accuracy, which is a defensible trade. The same count on a draft with thin evidence points at a draft that was never really worked on.

Recording the four rows

The sheet that comes out of this is small enough to keep. One row per dimension, each carrying its score out of ten and its weight, and a total that is the weighted mean of the four. Information gain carries 30 percent, specificity 25, source compliance 25, and argument depth 20.

Two habits keep the sheet useful. Record the four scores rather than the total alone, because the total is what a blended number hides and the four rows are what tell you where to spend the next hour. And record the evidence column next to each score, meaning the specific sentence or the specific missing source that produced it, so a later reader can disagree with the reading instead of only with the number.

The sheet is a record rather than a verdict. A draft at 7.8 with a 9.0 on source compliance and a 7.0 on argument depth is a different object from a draft at 7.8 built from four even rows, and the sheet is what preserves that difference after the numbers have been rounded.

What no check can tell you

None of the four lanes answers whether the piece should exist. Evaluation measures execution against a standard, and a draft can satisfy every row while arguing a point the audience has already settled, or restating a conclusion more confidently than its sources support.

There is also no instrument that settles an individual authorship question. A false-positive rate describes a population, and re-running a detector produces a second measurement of the same kind rather than a verification. When authorship matters, the evidence is provenance: drafts, revision history, and the sources you worked from.

Treat the lanes as answers to distinct questions, each with its own instrument and its own blind spot. Detection covers the tool's confidence, quality covers the draft against a publish standard, evidence covers each claim, and tone covers how the prose reads. Sorting a draft into these four and running them in order is the whole method.

The four-dimension diagnostic behind these figures is free to run on any draft, and the ForgeRank published dataset is public and re-runnable, so the numbers above can be checked rather than taken on trust.

Working on a draft right now? You can run any piece through the same 4-dimension quality read before it ships. It is free, no signup, at forgerankai.com.