How Phrasly Ultra Passes the AI Detectors Other Humanizers Can't?

Muhammad Usman Ali

14 min read

Pangram was the hardest public test in Phrasly's September 2026 benchmark. Here is how Ultra scored. How the comparison was run, and what the numbers actually mean.

AI detectors do not all evaluate or report text the same way. Each one uses its own model, thresholds, and score format. One passage can look human on one checker and AI-generated on another.

That is why Pangram 4 is the main stress test in this comparison. Pangram built it to catch AI-assisted writing and mixed AI-human text. Including edited drafts. So, how Phrasly Ultra passes AI detectors that trip up other tools comes down to one dated result.

In Phrasly's September 15, 2026 benchmark, the Phrasly Ultra AI humanizer earned a 99.8% average Pangram human score across 72 texts. 70 of 72 outputs were 100% human.

The three other humanizers averaged 50% or less. Phrasly ran and published the test, and the results describe this test set. Not a guarantee for every text or future detector version.

Below, we cover the method, the comparison, and the limits of the Phrasly Ultra AI detector’s data.


See the Model Behind the 99.8% Pangram Human Score.

Watermark-free rewrites, tested on the same 72 texts as three rival humanizers.


AI Detectors Do Not All Look for the Same Signals

Myth versus fact cards about how AI detectors evaluate text

AI detectors are classification systems. Not interchangeable truth machines. Each is trained on different data, uses its own architecture and thresholds. Each accepts different kinds of text, and reports results in its own format.

That is why the same passage can score differently across tools. And why no percentage transfers cleanly from one detector to another.

Three ideas get blurred all the time, and AI detection estimates whether text looks machine-generated. Plagiarism or similarity checking looks for overlap with existing sources.

Proof of authorship is something neither one can deliver on its own. Every detector also carries two kinds of error risk. A false positive (human writing flagged as AI) and a false negative (AI writing labeled human).

The major tools describe their systems very differently.

Turnitin scores only "qualifying text,”. Meaning long-form prose, and warns that it can misidentify human, AI-generated, and AI-paraphrased writing.

According to Turnitin's AI writing report guidance, results below 20% show an asterisk instead of a number.

That range carries a higher rate of false positives. That is Turnitin's own reporting rule. Not an industry-wide threshold.

One more correction worth making. It is outdated to say every detector relies only on perplexity and burstiness. Modern systems, including Pangram 4 and GPTZero, use trained deep-learning classifiers.

For a fuller primer, see how AI detectors work.

Detector

Publicly documented output or scope

Why scores should not be equated

Pangram 4

API returns fraction_human; classifies human, AI, and AI-assisted text

Its score definition differs from every other platform

Turnitin

Percentage of qualifying long-form prose likely AI-generated or AI-paraphrased

Qualifying-text and display rules, including the below-20% asterisk

GPTZero

Document-level and sentence-level classifications with confidence categories

Its own model and confidence interpretation

ZeroGPT / Originality.ai

Platform-specific classifications

Their percentages do not translate into a Pangram human score

Why Pangram Is the Benchmark to Watch

Big stat card showing Pangram 4 AUROC, false positive and false negative rates

Pangram matters here. Pangram 4 was built to identify fully AI-generated, AI-assisted, and mixed human-AI text. This includes edited content that can slip past simpler detectors.

Its authors explicitly designed for robustness against humanized writing.

A rewrite that holds up on Pangram has cleared a detector aimed squarely at the problem humanizers try to solve.

What the Pangram 4 Technical Report Claims

Pangram Labs and University of Maryland researchers released the Pangram 4 technical report on July 29, 2026 (arXiv:2607.27183).

These are the authors' figures. Measured on their own evaluation sets:

  • AUROC of 0.9916. A measure of how well the model separates human from AI text overall.

  • False-positive rate of 0.0041%. The report puts it at roughly 1 in 24,000 human texts wrongly flagged.

  • False-negative rate of 0.3396% for AI text the model misses.

  • Claimed gains in fine-grained edit detection. Mixed AI-human text boundary detection, out-of-distribution performance, and robustness to adversarial attacks.

Treat these as Pangram's reported results. Not universal real-world guarantees. Outside reviewers have noted that lab numbers can drop on short, mixed, or style-imitating text.

How the Pangram Human Score Works

Pangram's API returns fraction_human. The share of a text classified as human after its AI and AI-assisted categories are excluded. Phrasly multiplied that fraction by 100 to get a Pangram human score. Then averaged it across each tool's completed outputs.

A 100% human output means fraction_human was exactly 1 to four decimal places. The table below shows why a human score and an AI detector pass rate are different numbers.

Metric

What it measures

Phrasly Ultra, Sept 2026

Average Pangram human score

Mean of every output's human score (0 to 100)

99.8% across 72 outputs

Outputs scored 100% human

Outputs with fraction_human of exactly 1

70 of 72 (97.2%)

Pass rate at a threshold

Share of outputs above a chosen cutoff; Phrasly used at least 50% human in its competitor notes

72 of 72 at the 50% cutoff, per Phrasly's benchmark page

Remember that a human score is a detector assessment. It is not proof that a person wrote the text.

For accuracy, pricing, and features, read this breakdown of how Pangram AI detection works.

Is Pangram harder for humanizers? Pangram's own research emphasizes robustness to humanized and edited AI text.

In Phrasly's September 2026 comparison, the three non-Phrasly humanizers averaged between 13.3% and 50.0% on the Pangram human score. Phrasly Ultra averaged 99.8%.

What the Phrasly Ultra Benchmark Actually Tested

Phrasly gave four humanizers the same 72 long English AI-generated texts. Scored every rewrite with Pangram's pangram-4 model. And counted unsupported factual changes.

It then published the inputs, scores, and SHA-256 output fingerprints. Anyone can check the math. The goal was a like-for-like test with no hidden settings or cherry-picked samples.

Here is the method, step by step:

  • Phrasly randomly sampled 72 long English texts labeled "AI" in the public AI Text Detection Pile on Hugging Face (MIT license).

  • Inputs ran 604 to 900 words, averaging 761 words. For 54,800 words per tool and 219,200 words across all four tools.

  • Phrasly Ultra, StealthGPT, WriteHuman, and Undetectable AI each received identical inputs. With one output kept per tool per text.

  • Every completed output went through Pangram's AI detection API using the pangram-4 model. With fraction_human recorded to four decimal places.

  • Each output was compared with its input and checked for hallucination. Defined as any unsupported factual change or addition.

  • Per-tool CSVs list inputs, scores, word counts, notes, and output fingerprints. Output text is withheld. Shared only with approved researchers, journalists, and tested providers.

Phrasly AI evaluated and published this benchmark. Phrasly also owns Ultra. That is exactly why the records and formulas are public. The averages can be recomputed from the CSV files. The test also has clear limits:

  • It covers long English AI-generated texts only.

  • It reflects one benchmark edition and one Pangram model snapshot.

  • It is a vendor-run test. Even though the data is supplied for checking.

  • It does not establish how every genre, language, short passage, future detector version, or manually edited draft will score.

For every column, formula, and download, see the full Phrasly Ultra benchmark methodology and data.

Phrasly Ultra September 2026 Pangram benchmark methodology using 72 identical AI-generated texts

The Result: Ultra Averaged 99.8% Human on Pangram

In Phrasly's September 2026 benchmark, Phrasly Ultra averaged a 99.8% Pangram human score across 72 outputs. And 70 of them were rated 100% human. The next-best humanizer in the same test, StealthGPT, averaged 50.0%.

Ultra also recorded the fewest unsupported factual changes. Leading all four reported measures in this edition.

Humanizer

Avg. Pangram human score

Outputs scored 100% human

Avg. unsupported factual changes

Hallucination-free outputs

Phrasly Ultra

99.8% across 72

70 of 72 (97.2%)

0.85 across 72

35 of 72 (48.6%)

StealthGPT

50.0% across 72

19 of 72 (26.4%)

3.54 across 71

4 of 71 (5.6%)

Undetectable AI

24.5% across 72

3 of 72 (4.2%)

6.16 across 70

1 of 70 (1.4%)

WriteHuman

13.3% across 69

1 of 69 (1.4%)

2.64 across 69

10 of 69 (14.5%)

Source: Phrasly AI, September 2026 benchmark; last tested September 15, 2026; Pangram model pangram-4. Higher human scores are better, and fewer unsupported factual changes are better. WriteHuman declined 3 of the 72 inputs. Its averages cover 69 completed outputs. One StealthGPT output and two Undetectable AI outputs had no recorded hallucination count. So only their hallucination measures exclude those rows.

Phrasly Ultra Average Pangram human score compared with StealthGPT, Undetectable AI, and WriteHuman

So, what does this mean for Phrasly Ultra vs. other humanizers? The gap is large and consistent.

👊 Ultra averaged 99.8% human on Pangram. While the next-best humanizer in the same test averaged 50.0%. Keep the scope honest, though. Only three competing products were tested.

This does not show that every other humanizer falls short on Pangram. Anyone asking for the best AI humanizer for Pangram should read this as strong evidence from one controlled test. Not a universal ranking.

This roundup of the best AI humanizer tools covers the wider field.

The Phrasly Ultra Pangram score should also always be stated precisely. 99.8% is the average human score.  Not the percentage of tests passed. And the factual-accuracy result deserves equal attention.

Ultra performed best. But only 35 of 72 outputs were hallucination-free. Every rewrite still needs human fact-checking.


Want the numbers behind the headline? See every score, the method, and the downloadable data before you decide.


How Phrasly Ultra Is Built to Balance Detector Performance and Meaning

Phrasly says Ultra combines model fine-tuning and detector-informed reinforcement learning. And evaluation of meaning, grammar, and natural flow. Detector feedback is one training signal. Not the only quality measure.

The goal is a rewrite that reads naturally. Keeps the original meaning intact, rather than one tuned only to move a score.

Here is what Phrasly documents publicly about the Phrasly Ultra AI humanizer:

  • Phrasly starts from an existing foundation model. And adapts it for humanization. It does not claim to train a foundation model from scratch.

  • Training data, fine-tuning, the reinforcement-learning approach, and evaluation methods are developed in-house.

  • Detector feedback sits alongside checks for meaning preservation, natural flow and grammar. And overall writing quality.

  • Production humanization runs on models Phrasly deploys and manages. Including fallback routes. Phrasly says it does not send humanization text to OpenAI, Anthropic, Google, or another provider's LLM API.

And that its rewrites are watermark-free output by default because it controls the deployed models.

What Phrasly does not publish is a list of the linguistic features Ultra changes for any specific detector. This article does not speculate about token probabilities, sentence length, or detector weaknesses.

For the published technical detail, read how Phrasly trains and evaluates its humanization models.

Phrasly Ultra model selector with Balanced mode selected

What About GPTZero, ZeroGPT, Turnitin, and Originality.ai?

Phrasly's live Ultra page reports pass rates above 99% on GPTZero, ZeroGPT, Turnitin, and Originality.ai. It recommends Balanced mode as the starting point.

Those are Phrasly product claims. The public Pangram benchmark carries far more method detail than the landing page currently gives for those four detectors. The two kinds of evidence should be weighed differently.

So, what AI detectors does Phrasly Ultra pass? The honest answer comes in two tiers. One measured result with a full method. And four vendor-reported figures. Here is how each should be read:

Detector

Current Phrasly statement

Evidence type and how to read it

Pangram 4

99.8% average human score across 72 texts; 70 of 72 at 100% human

✅ Measured result with date, sample, model, owner, and limitations

GPTZero

Over 99% pass rate

Phrasly reports this; results vary; not comparable to Pangram's 99.8% score

ZeroGPT

Over 99% pass rate

Phrasly reports this; same attribution rule

Turnitin

Over 99% pass rate

Phrasly reports this; ❌ never a promise about a specific submission

Originality.ai

Over 99% pass rate

Phrasly reports this; same attribution rule

A "pass rate" and an average human score are not the same metric. A 99% claim on one tool cannot be stacked against 99.8% on another.

If you want background on the individual tools, see these reviews of the GPTZero AI detector and Turnitin AI detection. Plus this hands-on answer to: Does ZeroGPT work?

One distinction matters most. Turnitin's own guidance says its AI-writing percentage should not be the sole basis for an adverse decision. That is not a knock on detector companies. It is a reminder that every score needs human interpretation.

What the 99.8% Result Does Not Guarantee

No AI humanizer can guarantee the same result for every text or future detector version. A detector score does not prove who wrote a document. Or whether its author followed a policy.

Phrasly's result describes one dated test set of long English texts. Run and published by Phrasly. It should be read with those boundaries attached.

Here is the clean split between what the benchmark shows and what it cannot:

✅ What this benchmark shows

❌ What it does not show

Ultra led four measures on 72 identical long English texts

How every genre, short passage, or language will score

A 99.8% average Pangram human score on pangram-4

How future Pangram versions or other detectors will behave

Fewer unsupported factual changes than three rivals

Zero hallucinations or perfect meaning preservation

A dated, inspectable vendor-run comparison

An independent result, or proof of authorship

Detector behavior can shift with text length, genre, language, model version, amount of editing, and surrounding context. Phrasly also notes Ultra has not been officially or fully tested across every language.

The test measured detector classification and unsupported factual changes. Not every dimension of writing quality. So Ultra still needs human review for facts, citations, names, dates, claims, and voice.

Detectors make mistakes as well. For a closer look at how often, read can AI detectors be wrong.

If a third-party test of "Phrasly" shows lower results, check the test date and model first. HumanizerBench's September 2026 Phrasly cycle was last tested on September 2. Thirteen days before Ultra launched on September 15.

It should not be described as evidence for or against Ultra unless the test names the Ultra model.

Finally, use Ultra to refine writing you are permitted to revise. Keep your own ideas, and reduce robotic phrasing. Follow workplace, publisher, or institutional rules. And disclose AI assistance when required.

Phrasly’s full position is in its responsible-use policy.

How to Try Phrasly Ultra Without Losing Your Meaning

The best way to use Ultra is to treat it as an editing partner. Not a final step. Start on the recommended mode. Compare the rewrite with your source line by line. And restore anything that makes the draft yours.

Detector scores then become one quick check. Not the goal.

  • Open the Phrasly Humanizer and select Phrasly Ultra.

  • Start with Balanced, which the live Ultra page recommends for most text.

  • Compare the rewrite with the source for altered facts, missing qualifiers, changed citations, or a shift in tone.

  • Restore personal examples and phrasing that make the draft genuinely yours.

  • Treat any detector result as a screening signal. Not proof of authorship or quality.

  • Keep drafts or revision history when authorship matters. Follow the relevant AI-use policy.

Ultra offers five rewrite approaches, running from Gentle to Max. Balanced is the default. If Balanced reads too close to the original, step toward Max. Then re-check the facts. Since stronger rewrites change more text.


Start with Ultra's meaning-first Balanced mode and see how your own draft reads.


Phrasly Ultra AI Detectors Results at a Glance

Here is what to take away from the Phrasly Ultra benchmark:

  • Detectors differ in models, rules, and outputs. So cross-detector scores always need context.

  • Pangram was the most transparent, fully documented test in Phrasly's launch evidence.

  • 👊 Phrasly Ultra averaged 99.8% human and led the four-tool comparison on every reported measure.

  • The result is evidence for a defined test set. Not a guarantee, and 35 of 72 outputs were hallucination-free.

If you are looking for a humanizer that passes AI detectors, weigh detector performance alongside meaning, factual accuracy, and natural writing, then review every output yourself.

Frequently Asked Questions

Is 99.8% a pass rate?

No! The 99.8% figure is the average Pangram human score across 72 outputs in Phrasly's September 2026 benchmark. Not the share of tests passed. Separately, 70 of 72 outputs scored exactly 100% human.

A pass rate depends on a chosen cutoff. This is why the two numbers should never be swapped. See how much AI detection is acceptable for more on thresholds.

Is Phrasly Ultra free?

Ultra is included with every paid Phrasly plan and with the trial. Free accounts use the standard Balanced model rather than Ultra. Plan details and current terms are listed on the Phrasly pricing page.

This is the best place to confirm what your account includes before you start.

Does Phrasly Ultra add an AI watermark?

No, Phrasly Ultra does not add an AI watermark. Phrasly says its humanization models and fallback routes are watermark-free by default. Phrasly deploys and manages the generation models itself.

It does not route humanization text through another provider's LLM API.

Will Phrasly Ultra always pass Turnitin?

No tool can promise that. Phrasly reports a pass rate above 99% on Turnitin in its current product claim. But detector results vary by text, and institutional policies differ. Turnitin itself says its percentage should not be the sole basis for an adverse decision.

No one should expect a guaranteed submission outcome.

Is using an AI humanizer allowed for school or work?

It depends on the context and the policy, and using a humanizer to misrepresent AI work where it is prohibited is not acceptable. Using one to improve writing you are allowed to revise can be fine. Follow your institution's, employer's, or publisher's rules.

 Disclose AI assistance whenever it is required.

Is a 100% human detector score proof that a person wrote the text?

No! A 100% human score is a classifier result. Not proof of authorship. Detectors can mislabel human writing as AI and the reverse. This is why this guide to AI detector false positives recommends treating every score as one signal among several.

Written by

Muhammad Usman Ali

Pakistan

Muhammad Usman Ali is an experienced SEO content writer with 3+ years of professional writing experience. He specializes in AI tools, AI detection technologies, and search engine optimized content.

Share this article