AI humanizer benchmarks

Last tested:

What is the best AI humanizer?

In the September 2026 benchmark, Phrasly Ultra led all four measures across 4 AI humanizers and the same 72 texts: a 99.8% average Pangram human score, 70 of 72 outputs rated 100% human, and the fewest hallucinations (0.85 per output). All results are publicly verifiable.

Last tested
Inputs tested
72
Detector
Pangram (pangram-4)

Swipe to see every score

Results from the September 2026 benchmark: 4 AI humanizers rewrote the same 72 texts. Every output was scored with Pangram, and outputs were checked for hallucinations.
AI humanizerPangram detectionFactual accuracy
Avg. human scoreHigher is better100% human outputsHigher is betterAvg. hallucinationsLower is betterHallucination-free outputsHigher is better
Phrasly Ultra
99.8%
Across 72 outputs
97.2%
70 of 72 outputs
0.85per output
Across 72 outputs
48.6%
35 of 72 outputs
StealthGPT
50.0%
Across 72 outputs
26.4%
19 of 72 outputs
3.54per output
Across 71 outputs
5.6%
4 of 71 outputs
WriteHuman
13.3%
Across 69 outputs
1.4%
1 of 69 outputs
2.64per output
Across 69 outputs
14.5%
10 of 69 outputs
Undetectable AI
24.5%
Across 72 outputs
4.2%
3 of 72 outputs
6.16per output
Across 70 outputs
1.4%
1 of 70 outputs

Higher Pangram human scores and higher shares of 100% human and hallucination-free outputs are better. Fewer hallucinations per output is better. Hallucination counts were not recorded for 1 StealthGPT output and 2 Undetectable AI outputs. Those outputs are excluded from the hallucination measures only.

The results, explained

How each humanizer compares.

A closer look at each tool’s Pangram results and hallucinations in the September 2026 benchmark, set against what each provider advertises.

Phrasly Ultra

Phrasly Ultra benchmark results

Phrasly Ultra averaged a 99.8% Pangram human score across 72 outputs, and 97.2% of them (70 of 72) were rated 100% human. It averaged 0.85 hallucinations per output, and 48.6% of outputs (35 of 72) were hallucination-free.

Phrasly Ultra recorded the best result on the average Pangram human score, 100% human outputs, average hallucinations, and hallucination-free outputs among the 4 humanizers tested.

Phrasly AI ran this benchmark and scored Phrasly Ultra with the same Pangram model and hallucination definition as every other tool, under the same output publishing rule. The paid Unlimited plan includes unlimited humanization after the trial, with up to 5,000 words per request.

StealthGPT

StealthGPT compared with Phrasly Ultra

In our September 2026 test of 72 texts, StealthGPT’s Super model passed Pangram on 53% of texts (38 of 72), meaning Pangram rated the rewrite at least 50% human. Its average Pangram human score was 50.0%, and only 19 of its 72 rewrites received a fully human rating. Its rewrites averaged 3.54 hallucinations each, 4.2 times as many as Phrasly Ultra, and 4 of 71 had none.

StealthGPT’s API page lists its Super model at an 89% bypass rate against Pangram V4. A September 8, 2026 post by StealthGPT’s founder, Jozef Gherman, reports an internal test of 100 samples in which Pangram V4 classified 89 of them as human, and describes Super as a multi-stage pipeline that inspects, fixes, and re-inspects text, including checks for factual consistency. That 89% comes from StealthGPT’s own samples, counting outputs Pangram called human. On our 72 texts, counting every rewrite Pangram rated at least 50% human, Super passed on 53%.

Phrasly Ultra, a single-pass model with no multi-stage pipeline, passed Pangram on all 72 texts in the same test, averaged a 99.8% Pangram human score, and had 0.85 hallucinations per rewrite.

WriteHuman

WriteHuman compared with Phrasly Ultra

WriteHuman’s homepage reports a 94% Pangram human score after WriteHuman processing, and its site banner says the model was updated for Pangram on September 3, 2026. In our test on September 15, 2026, WriteHuman averaged a 13.3% Pangram human score across the 69 texts it rewrote, 80.7 percentage points below the 94% its homepage reports.

WriteHuman passed Pangram on 9% of the texts it rewrote (6 of 69), meaning Pangram rated the rewrite at least 50% human. Only 1 of its 69 rewrites received a fully human rating, it declined to rewrite 3 of the 72 texts, and its rewrites averaged 2.64 hallucinations each, 3.1 times as many as Phrasly Ultra.

WriteHuman’s own Pangram review says “no humanizer guarantees a specific score from Pangram or anyone else,” and the 94% figure on its homepage carries no test date, sample size, or method. Every number here comes with the texts we used, a score for each rewrite, and a fingerprint of every output, so anyone can check it.

Undetectable AI

Undetectable AI compared with Phrasly Ultra

Undetectable AI averaged a 24.5% Pangram human score across 72 outputs, and 4.2% of them (3 of 72) were rated 100% human. It averaged 6.16 hallucinations per output, and 1.4% of outputs with a hallucination count (1 of 70) were hallucination-free. Hallucination counts were not recorded for 2 of its outputs, so its hallucination measures cover 70 outputs.

Undetectable AI’s average Pangram human score was 75.3 percentage points lower than Phrasly Ultra’s. Undetectable AI averaged 5.31 more hallucinations per output than Phrasly Ultra.

Undetectable AI’s paid humanizer uses monthly word credits. Its separately marketed free unlimited tool should be distinguished from the paid model when comparing allowances and results with Phrasly AI.

Beyond the scores

The features behind the subscription.

Compare what each provider advertises alongside the measured results, including Pangram claims and humanization allowances.

Swipe to compare provider features

Advertised features confirmed by Phrasly AI for its own product and reviewed on competitor websites. Feature claims are separate from measured benchmark results.
ProviderPhrasly UltraStealthGPTWriteHumanUndetectable AI
Advertises Pangram bypassAdvertisedAdvertisedAdvertisedNot stated
Paid humanizationUnlimitedMixed wordingUnlimited on UltraMonthly credits
Published allowance

Unlimited humanization after the trial. Up to 5,000 words per request; fair-use terms apply.

Plans list word quotas. The pricing FAQ also advertises unlimited monthly output.

Basic and Pro list 80 or 200 humanizer requests per month. Ultra lists unlimited requests, with up to 3,000 words per request.

Paid humanizer plans use monthly word credits. A separate free tool is marketed as unlimited.

A checkmark means the provider advertises that feature; it is not verification that the feature works. Phrasly AI confirms its own features; competitor claims reflect the pages reviewed. Provider pages reviewed September 14 and 15, 2026.

Reading the results

What makes a good AI humanizer?

A good AI humanizer should produce text that detectors classify as human without inventing or changing facts. This benchmark measures both: how Pangram classifies each output, and how many unsupported factual changes it contains. Plan allowances are compared separately.

Average Pangram human score

Higher is better

Pangram’s AI detection API returns fraction_human: the share of a text it classifies as human, excluding its AI and AI-assisted categories. The human score is that fraction multiplied by 100, averaged across each tool’s outputs.

Pangram API reference

Outputs rated 100% human by Pangram

Higher is better

The share of a tool’s outputs that Pangram classified as entirely human: a fraction_human of exactly 1, as recorded to four decimal places.

Pangram API reference

Average hallucinations per output

Lower is better

The average number of unsupported factual additions or changes in each output, relative to its input. It is a count per output, not a percentage.

Hallucination-free outputs

Higher is better

The share of a tool’s outputs with no unsupported factual additions or changes relative to the input, among outputs with a recorded hallucination count.

The Phrasly AI approach

Why choose Phrasly AI as your AI humanizer?

Phrasly AI develops specialized humanization models through extensive fine-tuning and reinforcement learning. Our proprietary work connects model training, writing evaluation, and the infrastructure that delivers each rewrite.

Explore Phrasly AI technology
Phrasly AIInside the model pipeline
  1. 01Fine-tune
  2. 02Refine with feedback
  3. 03Evaluate
  4. 04Deploy

Proprietary training. Specialized models. Phrasly AI inference.

Quality is part of training

Helper agents review writing quality and meaning preservation throughout training and internal evaluations. Those checks are not part of this benchmark.

A feedback-driven model pipeline

Detector-informed reinforcement learning sits alongside writing-quality evaluation, helping our team refine how the model rewrites text.

Our deployed models, end to end

Production humanization runs entirely on Phrasly AI’s deployed models, including fallback routes. Rewrites are not sent to another provider’s LLM API.

Benchmark methodology

How we ran the benchmark.

Every humanizer received the same texts, and every output was scored the same way. These are the steps behind the September 2026 results.

September 2026 edition
01

The same 72 texts

We sampled 72 long texts at random from those labeled “ai” in the AI Text Detection Pile. They run 604 to 900 words (average 761). Every humanizer received the same texts, and we kept one output per tool per text.

02

Scored with Pangram

Each output was submitted to Pangram’s AI detection API (model pangram-4). We recorded fraction_human, the share of text classified as human, to four decimal places.

03

Counted hallucinations

Outputs were checked against their inputs for unsupported factual additions or changes. We report the average count per output and the share of outputs with none, among outputs with a recorded count.

04

Published the records

Declined texts and missing hallucination counts were excluded only from the measures they affect. Each tool’s CSV lists every output’s scores and SHA-256 fingerprint. Outputs Pangram rated 50% human or higher were withheld, whichever tool produced them, and all other outputs were published in full.

Where the texts come from

The inputs were sampled at random from long texts labeled “ai” in the AI Text Detection Pile, a public Hugging Face dataset that labels each text as AI-generated or human-written. The dataset is published under the MIT license.

View the dataset on Hugging Face

Why 72 texts

Every humanizer received the same fixed set of 72 long texts, 604 to 900 words each (54,800 words per tool), so the comparison is fair. A set this size is also practical to run through every humanizer’s paid plan (219,200 input words across the 4 tools) while still giving each tool a good sample.

About this edition

Edition
September 2026 (humanizer-benchmark-2026-09)
Last tested
September 15, 2026
Inputs tested
72
Evaluation by
Phrasly AI
Detector
Pangram AI detection API, model pangram-4. The human score is fraction_human multiplied by 100, averaged across each tool’s completed outputs. An output counts as 100% human when fraction_human is exactly 1.
Hallucination count
The number of unsupported factual additions or changes in an output relative to its input. Lower is better. Hallucination measures cover outputs with a recorded count.
Refusals and missing counts
WriteHuman declined 3 of 72 texts. Declined texts returned no output and are excluded from all of that tool’s averages. Hallucination counts were not recorded for 1 StealthGPT output and 2 Undetectable AI outputs. Those outputs are excluded from the hallucination measures only.
Input set
72 AI-generated English texts of 604 to 900 words (average 761), 54,800 words per humanizer, sampled at random from long texts labeled “ai” in the AI Text Detection Pile (MIT license). Every tool received the same texts, with one output per tool per text.
Output publishing
One rule for every tool: outputs Pangram rated 50% human or higher were withheld from the public downloads, and all other outputs were published verbatim. Every completed output is listed with its scores and a SHA-256 fingerprint of its exact text.
Models and settings
Phrasly Ultra ran as a single pass, with no multi-stage pipeline. StealthGPT was tested with its Super model. Undetectable AI was run through its humanizer API with the Undetectable model v11sr, University readability, General Writing purpose, and Balanced strength. No other settings were recorded for any tool, including Phrasly Ultra’s own mode and WriteHuman’s model or mode.

Follow the product’s development

Our technology page explains the model pipeline. The product changelog records releases and improvements.

Product changelog

Explore the evidence

Download the records behind the results.

Each tool has its own CSV with every input, Pangram result, hallucination count, word count, and a SHA-256 fingerprint of each output. One rule applies to every tool: outputs Pangram rated 50% human or higher are withheld, and all other outputs are included in full. The input_source column credits the AI Text Detection Pile (MIT license) as the source of the inputs.

Phrasly Ultra

CSV

72 inputs and 72 scored outputs. All 72 outputs withheld (each rated 50% human or higher).

StealthGPT

CSV

72 inputs and 72 scored outputs. 34 outputs published, 38 withheld.

Outputs without a hallucination count: 1, marked in the notes.

WriteHuman

CSV

72 inputs and 69 scored outputs. 63 outputs published, 6 withheld.

3 declined texts, marked as refused with no output.

Undetectable AI

CSV

72 inputs and 72 scored outputs. 61 outputs published, 11 withheld.

Outputs without a hallucination count: 2, marked in the notes.

Our training data, model weights, and development recipes remain proprietary.

A little more detail

AI humanizer benchmark questions

How the benchmark was run, what the scores mean, and how to check the data yourself.

What is the best AI humanizer?

In the September 2026 benchmark, Phrasly Ultra led all four measures across 4 AI humanizers and the same 72 texts: a 99.8% average Pangram human score, 70 of 72 outputs rated 100% human, and the fewest hallucinations (0.85 per output). All results are publicly verifiable.

What does this AI humanizer benchmark compare?

The September 2026 edition compares Phrasly Ultra, StealthGPT, WriteHuman, and Undetectable AI on the same 72 long AI-generated English texts, with one output per tool per text. It reports four measures: the average Pangram human score, the share of outputs Pangram rated 100% human, average hallucinations per output, and the share of hallucination-free outputs.

Why did Phrasly AI run this benchmark?

Most humanizer comparisons are published by humanizer companies themselves, often without the inputs, settings, or outputs needed to check them. We wanted a test anyone can check: the same 72 texts for every tool, one output per tool per text, the same Pangram model and hallucination definition for everyone, and every input, score, and output fingerprint published. We also compared each competitor’s advertised claims with what we measured.

Where do the benchmark texts come from?

The 72 inputs were sampled at random from long texts labeled “ai” in the AI Text Detection Pile, a public Hugging Face dataset published under the MIT license that labels each text as AI-generated or human-written. They are AI-generated English texts of 604 to 900 words, and every humanizer received the same set. Each CSV credits the dataset in its input_source column.

Why does the benchmark use 72 texts?

A fixed set of 72 long texts keeps the comparison fair: every humanizer received the same texts, 604 to 900 words each (average 761), for 54,800 words per tool. A set this size is also practical to run through every humanizer’s paid plan (219,200 input words across the 4 tools) while giving a good sample for each tool. The results describe these texts and may differ on other kinds of writing.

What does the Pangram human score mean?

Pangram’s AI detection API (model pangram-4) returns fraction_human: the share of a text it classifies as human, excluding its AI and AI-assisted categories. The human score is that fraction multiplied by 100 and averaged across a tool’s outputs. Higher is better. It is a detector assessment, not proof of human authorship.

What does “100% human outputs” mean?

It is the share of a tool’s outputs that Pangram classified as entirely human, with a fraction_human of exactly 1 as recorded to four decimal places. An output with any portion classified as AI or AI-assisted does not count.

What counts as a hallucination?

A hallucination is an unsupported factual addition or change relative to the input text. The table reports the average count per output, where lower is better, and the share of outputs with none. It is a count, not a percentage or a score out of 100.

How were refusals and missing hallucination counts handled?

WriteHuman declined 3 of 72 texts. Declined texts returned no output and are excluded from all of that tool’s averages. Hallucination counts were not recorded for 1 StealthGPT output and 2 Undetectable AI outputs. Those outputs are excluded from the hallucination measures only. Each case is marked in the per-tool CSV files.

Does an advertised Pangram bypass mean a tested result?

No. The feature table records what each provider advertised on the pages reviewed, and a checkmark is not verification that a feature works. The benchmark table shows measured Pangram results on the same 72 texts, so you can compare each claim with the outcome.

Can I download the benchmark data?

Yes. Each tool has its own CSV with every input, its source dataset, the Pangram model, fraction_human, human score, hallucination count, word counts, a SHA-256 fingerprint of each output, and notes. Outputs Pangram rated 50% human or higher are withheld from every tool’s file, and all other outputs are included in full. In this edition, 72 of 72 Phrasly Ultra outputs, 38 of 72 StealthGPT outputs, 6 of 69 WriteHuman outputs, and 11 of 72 Undetectable AI outputs were withheld.

Why are some outputs not in the public download?

Detector developers improve their models by training on humanized text, and outputs that detectors currently classify as human are the most useful training examples. Pangram’s DAMAGE paper (arXiv:2501.03437), for example, describes a detector made robust to humanizer output through data augmentation. To keep passing outputs out of detector training sets, we withhold every output Pangram rated 50% human or higher, from every tool, including all of Phrasly Ultra’s outputs. Each withheld output is still listed with its scores and a SHA-256 fingerprint, so anyone who receives the texts can confirm they are exactly what we tested.

How can I request withheld outputs and verify them?

Researchers, journalists, and the providers we tested can use the request form on this page to describe their intended use. Access requires agreeing not to use the texts to train or tune AI detectors and not to republish them. The Phrasly AI team reviews every request and responds within 10 business days, and nothing is released automatically. To verify a text you receive, compute the SHA-256 hash of its exact UTF-8 text and compare it with output_sha256 in the CSV. A match confirms the text is the one that was scored.

Who publishes this benchmark, and how can I verify the results?

Phrasly AI ran and published this benchmark, which includes its own Phrasly Ultra. Scores are listed for every completed output, including withheld ones, so you can recompute each average from the CSV files: average the human score over completed outputs, count outputs with a fraction_human of exactly 1, and average the hallucination counts over outputs that have one. The fingerprints show that no output was swapped after scoring.

Research access

Request the withheld outputs.

Outputs Pangram rated 50% human or higher are withheld from the public downloads, whichever tool produced them. Researchers, journalists, and the providers we tested can request the full texts by describing how they plan to use them.

  • Open to researchers, journalists, and the providers tested
  • Texts may not be used to train or tune AI detectors, or republished
  • The Phrasly AI team reviews every request and responds within 10 business days
  • Nothing is released automatically
  • Check each text you receive against its SHA-256 fingerprint in the CSV

Request withheld benchmark outputs

Open to researchers, journalists, and the providers we tested. By requesting access, you agree not to use the texts to train or tune AI detectors and not to republish them. We review every request and reply within 10 business days.

We’ll use these details to review and respond to your request. Privacy policy

Put your own words to the test.

Experience Phrasly AI with the text that matters to you.

Try Phrasly AI