A humanizer rewrites surface phrasing while leaving structure intact. It can lower a detector's confidence score, because paraphrase attacks genuinely degrade classifier accuracy, according to Krishna et al. (arXiv 2303.13408). It cannot raise a draft's quality, because the ForgeRank published dataset shows phrasing cleanup moved no draft up a score band across 29 scored drafts.
Why This Matters
Open a draft that scored 4.0 on all four quality dimensions and run it through a paraphrasing tool. The phrasing gets cleaner. The score does not move. In the ForgeRank published dataset, ten drafts sat flat at 4.0 across every dimension, and ten others held 4.0 on three dimensions with source compliance at 7.0, landing at 4.8. Source attachment was the only dimension that moved a headline score on its own.
That gap between detection and quality is where most rewriting effort gets wasted. A writer who wants a better draft and a writer who wants a lower detection score are working on different problems, and the tools that solve one do nothing for the other.
We do not build a humanizer and we do not recommend one. This page is not a guide to evading detection. It explains what the rewrite changes, where it stops, and what to fix instead.
What Does a Humanizer Change at the Sentence Level?
A humanizer swaps content words while preserving syntax. Function words keep their positions and proportions, sentence boundaries stay where they were, and punctuation habits survive the rewrite. The output reads differently and scans identically.
Mitchell et al. (arXiv 2301.11305, ICML 2023) showed why this works on detectors. Machine-written text sits near a local optimum of the model's own probability, so masking a handful of tokens lowers its likelihood under that model. Human text stays roughly flat under the same perturbation. A paraphrase attack moves the text off that optimum, and Krishna et al. measured detector accuracy dropping hard as a result.
The defense that restored accuracy was not a stronger classifier. It was retrieval against a database of known machine-generated outputs. That detail matters for anyone treating a passing score as durable. The score moved because the surface moved, and the surface is the cheapest thing to change.
Can Rewriting Lower a Detection Score?
Yes, and the effect is measurable. Krishna et al. found that paraphrase attacks dropped detector accuracy substantially, which is why rewriting tools exist and why they sell. If your only goal is a different number from a detector, rewriting is the shortest path to it.
The instability sits on the other side of the same finding. OpenAI retired its own AI classifier in July 2023 after reporting it identified only 26 percent of AI-written text while flagging 9 percent of human-written text as AI. Liang et al. (arXiv 2304.02819, published in Patterns) found roughly 61 percent of TOEFL essays written by non-native English speakers were flagged as AI-generated, with no model involved in the writing.
A 61 percent false-positive rate on human writing means the score you are optimizing against is not measuring authorship. It is measuring how closely your prose resembles the detector's training distribution. Rewriting to satisfy that target means writing toward the distribution that produced the flags in the first place.
Does Cleaner Phrasing Make a Draft Better?
No. The ForgeRank published dataset tracked 29 scored drafts through a phrasing pass and recorded 594 hits across 23 distinct families, with em-dash density alone firing 182 times across 19 drafts. Cleaning all of that moved no draft up a score band. Eight of the nine drafts scoring 7.0 or above stayed flagged as heavy on AI phrasing after cleanup.
The assumption underneath the rewrite is that AI-sounding phrasing is what caps a draft's score. The distribution says otherwise. Drafts with clean phrasing and no attached sources still scored at the bottom, and the drafts that climbed did so on source compliance, not on sentence polish.
Here is the boundary. A paraphrase aimed at a detector is aimed at the wrong object. The specific detail that makes a draft worth reading, the named figure, the attached source, the argument that takes a position, is the first thing a rewrite strips, because specificity is the least compressible part of a sentence and the hardest to preserve through a word swap.
How Do You Make AI-Assisted Writing Read as Human?
Attach sources to claims, add detail only you can supply, and argue a position. Then score the draft instead of paraphrasing it. The published distribution rewards source attachment above every other dimension, and that is a change a rewrite cannot make for you.
Concrete levers, in the order they move a score:
- Attach a source to every figure. In the ForgeRank dataset, drafts that held source compliance at 7.0 scored 4.8 while identical drafts without it sat at 4.0. A number with a named source is checkable. A number without one is a liability.
- Replace generic claims with specific ones. "Many writers struggle with this" carries no information. "594 phrasing hits across 23 families" carries a fact a reader can verify.
- Take a position and defend it. Hedged prose reads as generated because it commits to nothing. A draft that states a stance and supports it reads as authored.
- Score before you rewrite. Run the draft against a rubric that measures source attachment, specificity, and argument. If the score is flat, the rewrite will not move it.
The ForgeRank methodology page describes how that scoring works. The point of scoring first is to find out whether your problem is phrasing or substance. Phrasing cleanup is cheap and cosmetic. Substance is what the distribution pays for.
What Most People Get Wrong About Humanizing AI Text
The mistake is treating detection and quality as the same target. They are not, and the tools built for one actively work against the other.
Google's March 2024 core update stated its aim as reducing low-quality, unoriginal content in search results by 40 percent, according to Google Search Central. That target is unoriginality, not authorship. A paraphrased draft with no sources and no argument is still unoriginal, and the rewrite made it harder to fix because it flattened the specific details that would have differentiated it.
The belief persists because detector scores are visible and quality scores are not. A writer gets a number back from a detector and treats it as the finish line. The number that predicts whether the draft gets read, cited, or ranked is the one nobody runs.
Key Takeaways
- A humanizer swaps content words while preserving syntax, so function words, sentence boundaries, and punctuation habits survive the rewrite.
- Krishna et al. (arXiv 2303.13408) measured paraphrase attacks dropping detector accuracy, so rewriting does move a detection score.
- The ForgeRank published dataset shows 594 phrasing hits across 23 families and 182 em-dash firings across 19 drafts, with no draft moving up a score band after cleanup.
- OpenAI's July 2023 classifier withdrawal reported 26 percent detection of AI text and 9 percent false positives on human text, and Liang et al. flagged 61 percent of human-written TOEFL essays.
- Source attachment was the only dimension in the ForgeRank dataset that moved a headline score on its own, lifting drafts from 4.0 to 4.8.
Frequently Asked Questions
Can a humanizer get past AI detectors?
It can lower a detector's confidence score, because paraphrase attacks degrade classifier accuracy, per Krishna et al. It cannot guarantee a pass against any named detector, and the target shifts as detectors retrain.
Why do AI detectors flag human writing?
Detectors score how closely text resembles machine output, not who wrote it. Liang et al. found roughly 61 percent of TOEFL essays by non-native English speakers were flagged as AI-generated despite no model involvement.
Does rewriting AI text improve its quality?
No. The ForgeRank published dataset tracked 29 drafts through a phrasing pass and found no draft moved up a score band. Eight of nine drafts scoring 7.0 or above stayed flagged as heavy on AI phrasing after cleanup.
What actually makes AI-assisted writing read as human?
Attached sources, specific detail, and a defended argument. In the ForgeRank dataset, source compliance was the only dimension that moved a headline score alone, lifting drafts from 4.0 to 4.8.
Is using a humanizer against Google's guidelines?
Google's March 2024 core update targeted low-quality, unoriginal content, not AI authorship. A paraphrased draft with no sources and no argument stays unoriginal, so the rewrite does not address what the update penalizes.
How many phrasing issues does a typical AI draft contain?
The ForgeRank published dataset recorded 594 hits across 23 distinct families in 29 drafts, with em-dash density firing 182 times across 19 drafts. Phrasing issues cluster in a small number of recurring families.
The Bottom Line
Paraphrasing moves a detection score and leaves quality where it was. If your draft scored flat before the rewrite, it scores flat after, and you have spent the effort on the dimension that does not predict whether anyone reads the piece. Open the ForgeRank methodology page and run your current draft against the source-compliance dimension first. If a claim in it has no attached source, add one before you change a single sentence.