What Pangram Detects and Why Many AI Humanizers Fall Short

Muhammad Usman Ali

10 min read

Does Pangram detect humanized AI text? Often, yes. Pangram is built for it. It sorts writing into human, AI-assisted, and AI-generated passages. Then runs a separate humanizer check. So, swapping words rarely changes its verdict.

That is why so many people ask whether Pangram detects humanizers after a rewritten draft still gets flagged.

In Phrasly's September 15, 2026 benchmark, Phrasly Ultra averaged a 99.8% Pangram human score across 72 long texts. 70 were rated 100% human. The other three humanizers averaged 50% or less.

You'll see what Pangram looks for, why many rewrites fall short, and what independent research says. And also how to check your own results.

Disclosure: Phrasly ran this benchmark and makes Phrasly Ultra. Scores describe the tested outputs. Not a guarantee for every text.


Humanize Your Draft with Phrasly Ultra

Paste your draft. Choose Phrasly Ultra and compare the rewrite with your original.


What Does Pangram Detect?

Pangram is an AI detector from Pangram Labs. Pangram labels each passage of a document as human, AI-assisted, or AI-generated. Then gives the whole document a Human, Mixed, or AI label.

Pangram 4 is released on July 29, 2026. It also runs a separate check for signs of a humanizer. It needs at least 50 words of prose.

How does Pangram detect AI? It reads a document in passages and classifies each one. Human means a person wrote it. AI-Assisted means a person and AI both shaped it, and AI-Generated means AI wrote it.

Pangram is built for long-form prose. Not short replies, code, or templated text. It does not check a list of "AI words". Its score does not prove who wrote something.

What is a good Pangram score?

According to the Pangram 4 model card, a document is labeled Human when at least 90% of its characters are classified as human.


Text is classified as AI when at least 80% of its characters are classified as AI-generated. Anything in between is a Pangram Mixed result. This often reflects Pangram AI-assisted passages.

The humanizer flag works separately. Pangram gives a humanizer score from 0 to 1. Flags text as humanized at 0.91 or above by default.

Pangram output

What it tells you

Passage labels (Human, AI-Assisted, AI-Generated)

Which parts of the text read as human or AI

Document label (Human 90%+, AI 80%+, otherwise Mixed)

The overall verdict a reader sees first

fraction human

The share of the text classified as human. This is the number Phrasly's benchmark reports.

Humanizer flag (score 0 to 1, flagged at 0.91+)

A separate signal that a humanizer may have processed the text

That is how Pangram works in brief.

For features and accuracy, see this Pangram AI detector review, or read how AI detectors work for the basics.

Does Pangram Detect Humanized AI Text?

Yes, often. Pangram trains its detector on humanizer output. And runs a dedicated humanizer check.

In its Pangram 4 technical overview, Pangram reports finding AI involvement in 98.83% of texts from 13 commercial humanizers.

That is the company's own test. Results still vary by tool, text, and version.

Pangram humanizer detection works because the detector has learned what humanizers do.

In the DAMAGE paper, Pangram's researchers describe training on humanized and paraphrased text on purpose, using data augmentation.

In August 2025, Pangram also described the traces many humanizers leave. Odd synonyms; it calls them "tortured phrases". Unnatural spacing, the same replacement phrase repeated, and non-standard characters.

Each one is a pattern a model can learn. Pangram does not claim 100% accuracy on humanized text. That is why tool-by-tool tests, like the benchmark below, matter more than any single headline number.

Why Do Many AI Humanizers Fall Short on Pangram?

Workflow diagram of four reasons AI humanizers fail Pangram detection

Many humanizers change the words but not the pattern underneath. So Pangram still reads the text as AI. The tricks some tools use to hide that pattern, like odd synonyms and hidden characters, become new signals.

Pushing harder to avoid detection often damages the writing itself. Seeing humanized text still detected as AI usually comes down to one of four reasons. Together, they explain why humanizers fail Pangram.

Swapping Words Does Not Change the Pattern

Synonym swaps keep the sentence plan, rhythm, and structure. This pattern Pangram learned from AI writing. The surface changes. But the order of ideas and the shape of each sentence stay the same.

So, can Pangram detect paraphrased text? In many cases, yes. Simply rewording the content may not be enough to change the underlying patterns that it looks for.

This guide to AI humanizer vs paraphrasing tools explains the difference.

Humanizer Tricks Become New Signals

Tortured phrases, spacing tricks, and unusual characters were meant to confuse detectors. Pangram has said it now looks for exactly these traces.

A trick that once helped can become the reason an AI humanizer still flagged your draft.

Dodging the Detector Can Break the Writing

Aggressive rewriting also changes meaning. In Phrasly's benchmark, the three competing tools averaged 2.64 to 6.16 unsupported factual changes per output. The made-up examples below show what that looks like.

Original

Weak rewrite

What went wrong

The trial included 240 participants.

The trial involved hundreds of volunteers.

❌ Lost the exact number

The policy may reduce delays.

The policy eliminates delays.

❌ Turned a maybe into a promise

"We found no effect," the authors wrote.

"There was a major improvement," the authors explained.

❌ Reversed a quote

Advertised Pangram Scores Do Not Always Hold Up

StealthGPT advertises an 89% pass rate on Pangram V4 from its own 100-sample test. In Phrasly's test, it passed 38 of 72.

WriteHuman’s homepage reports a 94% Pangram human score with no date or sample size. It averaged 13.3% in Phrasly's test. Samples, cutoffs, and dates differ. Treat any vendor's number as a claim to check.

This piece on why AI humanizers don't work covers the wider problem.

What Does Independent Research Say About Pangram and Humanizers?

Independent research backs Pangram in some settings. And shows gaps in others. A June 2026 peer-reviewed study found Pangram the strongest of four detectors on texts that included humanized AI.

An August 2026 study found that most humanized research abstracts got past an older Pangram version. Neither study tested Phrasly Ultra.

In "Who wrote this?", published in the International Journal for Educational Integrity on June 29, 2026, Van Vlasselaer, Van Droogenbroeck, and Spruyt tested GPTZero, Pangram, and Copyleaks, and Turnitin on 160 documents of known authorship. Including humanized AI text. Pangram performed best of the four tools overall. The study predates Pangram 4.

In an August 2026 study from the University of Notre Dame, Karr Jr., Khvatskii, Hua, and Chawla had Undetectable AI v11 rewrite 642 research abstracts.

Fewer than 4% of the AI rewrites stayed flagged by Pangram 3.2 or GPTZero at a 0.50 threshold. That result covers an older Pangram version, research abstracts only, and proxy labels.

The studies differ because they used different Pangram versions, text types, humanizers, and cutoffs.

A detector score is one signal. Not proof, as this guide explains, and AI detectors can be wrong.

Which AI Humanizer Passes Pangram?

In Phrasly's September 2026 benchmark of four humanizers, Phrasly Ultra was the only one to pass Pangram on all 72 texts. It averaged a 99.8% Pangram human score. 70 of 72 outputs were rated 100% human.

StealthGPT passed 38 of 72 texts. WriteHuman passed 6 of the 69 texts it rewrote.

If you are looking for the best humanizer for Pangram, keep that result scoped to this one disclosed test. Phrasly AI ran it. The last testing was on September 15, 2026, with Pangram's API (model pangram-4).

All four tools rewrote the same 72 AI-generated English texts. 604 to 900 words each (average 761). Randomly sampled from the public AI Text Detection Pile dataset. Each tool produced one output per text.

Humanizer

Outputs scored

Avg. Pangram human score

100% human outputs

Passed (50%+ human)

Phrasly Ultra

72

99.8%

70 of 72 (97.2%)

✅ 72 of 72

Stealth GPT (Super model)

72

50.0%

19 of 72 (26.4%)

38 of 72 (53%)

Undetectable AI

72

24.5%

3 of 72 (4.2%)

Not published

Write Human

69 (declined 3)

13.3%

1 of 69 (1.4%)

6 of 69 (9%)

Tested September 15, 2026 with Pangram's API (pangram-4). n=72 per tool, n=69 for WriteHuman. Ultra ran as a single pass. Mode not recorded. StealthGPT used Super. Undetectable AI used v11sr, University readability, General Writing, Balanced strength. WriteHuman's model was not recorded.

📊 How to Read These Numbers?

99.8% is an average. Not a pass rate. It is Ultra's average Pangram human score (fraction human) across 72 outputs.

A pass means at least 50% human, Phrasly's cutoff. Pangram's own Human label needs 90%. A fraction human of 1 means every character is read as human. So 70 of Ultra's 72 outputs met Pangram's human-share bar for a Human label.

Two things were not published. The benchmark did not report Pangram's separate humanizer flag for any output. It did not publish a pass count for Undetectable AI. So this doesn't state one.

For Phrasly Ultra Pangram results in context, see how Phrasly Ultra performs on other AI detectors.


Want to see how your own draft reads after a rewrite? Try the Phrasly Ultra humanizer model with the trial or any paid plan. Then check the result against your original.


Does a High Pangram Score Mean the Rewrite Is Accurate?

Quote card on Pangram scores not guaranteeing factual accuracy

No! A detector score only shows how the text reads to Pangram. Not whether your facts survived.

In the same benchmark, Phrasly Ultra had the fewest unsupported factual changes. 0.85 per output, yet 37 of its 72 rewrites still had at least one. Always check a rewrite against your original.

A hallucination means an unsupported factual addition or change compared with the input, counted per output. AI humanizer hallucinations matter because a humanizer changes meaning. The detector score will not warn you.

Humanizer

Avg. hallucinations per output

Hallucination-free outputs

Phrasly Ultra

0.85 (n=72)

✅ 35 of 72 (48.6%)

Stealth GPT

3.54 (n=71)

4 of 71 (5.6%)

Undetectable AI

6.16 (n=70)

1 of 70 (1.4%)

Write Human

2.64 (n=69)

10 of 69 (14.5%)

🔎 What to Check in Every Rewrite

Compare names, numbers, dates, quotes, citations, and any "may" or "might" claim with your original. Those are the details a rewrite most often drops or hardens.

To keep your own voice too, see how to make Ultra output sound like you.

What Does This Benchmark Not Prove?

The benchmark shows how four humanizers scored on one Pangram version. One set of long English AI texts and one date. Phrasly ran it and makes Ultra. It does not prove results for every text. Every detector or future Pangram versions.

  • Who ran it: Phrasly tested its own product against three competitors.

  • What was tested: 72 long English texts from a public dataset. Short copy, essays, or other languages may score differently.

  • When and which version: One date (September 15, 2026) and one Pangram model (pangram-4).

  • What was not reported: Pangram's humanizer flag and Ultra's mode. Output texts are withheld to keep them out of detector training. But researchers can request them and verify each one by its SHA-256 fingerprint.

How Do You Use Phrasly Ultra and Check the Result?

Save your original. Rewrite it with Phrasly Ultra starting in Balanced mode. Compare the result line by line before you use it. Ultra is a model inside the Phrasly AI Humanizer.

It is available with the trial and every paid plan. If you then run a detector, record which one, its version, and the date.

  • Save your original draft and note its key facts.

  • Open the Phrasly AI Humanizer and pick Phrasly Ultra in the model menu (trial or paid plan).

    This Phrasly Ultra vs Phrasly Humanizer comparison explains the difference.

  • Start with Balanced mode inside Ultra. Phrasly suggests Max only if a detector still flags the Balanced result. Balanced mode in Ultra is not the same as the standard Balanced model on free accounts.

  • Compare the rewrite with your original: Facts, quotes, numbers, and tone.

  • If you run a detector, note which one, its version, and the date.

    Phrasly's AI Detector gives 3 free checks. Then unlimited free use with a free account. It is Phrasly's own detector. Not Pangram.

  • Follow your school's or employer's AI rules. Disclose AI use where required.

Phrasly Ultra selected in the Phrasly AI Humanizer with Balanced mode

Ready to try it on your own writing? Rewrite and review your draft with Phrasly Ultra. Then read it once more before you submit.


Frequently Asked Questions

Does Phrasly Ultra Pass Pangram?

In Phrasly's September 2026 benchmark, yes. All 72 Ultra rewrites scored at least 50% human on pangram-4, Phrasly's pass cutoff.

Ultra averaged a 99.8% human score. 70 of 72 outputs were rated 100% human. Phrasly ran the test. Results can vary by text.

Can Pangram Be Wrong?

Yes! Pangram reports a 0.0041% false positive rate. About 1 in 24,000 human documents, on its own English benchmark. But outside tests vary.

Pangram false positives can still happen. A score is not proof. See why your essay is detected as AI.

Can Pangram Detect Paraphrasing Tools?

Often! Pangram trains on reworded and humanized text. So, paraphrasing alone rarely changes its verdict. Results still depend on the tool. The text and the Pangram version.

Does Watermark-Free Text Pass Pangram?

Not automatically. Phrasly says its models are watermark-free by default. But Pangram reads writing patterns and does not need a watermark. Watermark-free and passing a detector are separate things. Learn more about AI text watermarks.

Is Phrasly Ultra Free?

Ultra is included with the trial and every paid plan. Free accounts use the standard Balanced model instead.

What If Pangram Flags My Own Writing?

Keep your drafts, notes, and version history. Ask for a human review. A detector score alone does not show how a text was written.

Written by

Muhammad Usman Ali

Pakistan

Muhammad Usman Ali is an experienced SEO content writer with 3+ years of professional writing experience. He specializes in AI tools, AI detection technologies, and search engine optimized content.

Share this article