Frozen at publication
Review archive: September 2026
In the September 2026 HumanizerEval review, StealthGPT Super ranked first of 10 humanizers across 5 detectors. This was the first published comparison review.
Run details
- Review
- September 2026
- Run date
- September 30, 2026 (America/New_York)
- Methodology version
- v1.0
- Scoring script version
- v1.0
- Prompt set version
- None; production sample
- Humanizers
- StealthGPT Super; WriteHuman; Grubby; Rephrasy; Undetectable.ai; RewriteAI; HIX Bypass; Humbot; BypassGPT; Ryne
- Detectors
- Pangram 4.0; GPTZero 2026-09-13-base; Originality.ai turbo; ZeroGPT default; Winston AI 5.0
- Sample
- 50 English texts, 300 to 995 words, mean 596; same source texts for every humanizer
- Detector access
- Vendor APIs; the same output sent to every detector
- Report ID
- 2026-09-30-humanizer-comparison
- SHA-256 of scores.csv
- e78de0985b8c057da00293a0606548066299508e55bcd2a708e62526549b0f25
Results
| # | Humanizer | Pangram | GPTZero | Originality.ai | ZeroGPT | Winston AI | |||
|---|---|---|---|---|---|---|---|---|---|
| 1 | StealthGPTPassed all 5 detectors on 46 of 50 | 96.8% | 92% | 98% | 98% | 98% | 98% | 83.5% | 1.04 |
| 2 | WriteHumanPassed all 5 detectors on 32 of 49 | 92.7% | 65% | 100% | 98% | 100% | 100% | 58.7% | 1.02 |
| 3 | GrubbyPassed all 5 detectors on 18 of 50 | 85.6% | 36% | 94% | 98% | 100% | 100% | 77.5% | 1.11 |
| 4 | RephrasyPassed all 5 detectors on 20 of 50 | 84.0% | 46% | 88% | 86% | 100% | 100% | 54.5% | 1.13 |
| 5 | Undetectable.aiPassed all 5 detectors on 15 of 50 | 84.0% | 32% | 96% | 94% | 100% | 98% | 56.5% | 1.21 |
| 6 | RewriteAIPassed all 5 detectors on 13 of 50 | 77.2% | 28% | 86% | 88% | 88% | 96% | 72.0% | 1.09 |
| 7 | HIX BypassPassed all 5 detectors on 1 of 50 | 66.4% | 2% | 68% | 64% | 100% | 98% | 46.5% | 1.11 |
| 8 | HumbotPassed all 5 detectors on 0 of 50 | 66.4% | 0% | 62% | 70% | 100% | 100% | 47.5% | 1.12 |
| 9 | BypassGPTPassed all 5 detectors on 0 of 50 | 64.4% | 0% | 56% | 66% | 100% | 100% | 49.0% | 1.12 |
| 10 | RynePassed all 5 detectors on 0 of 50 | 23.6% | 22% | 8% | 0% | 62% | 26% | 79.0% | 1.07 |
Each column counts how many of the 50 texts that detector called human-written. Rank is the average of the five columns. WriteHuman returned 49 of 50 outputs, so its rates are out of 49. Full methodology.
- Pangram
- GPTZero
- Originality.ai
- ZeroGPT
- Winston AI
Pangram 4.0
| Humanizer | Strict bypass | Texts passed | Mean human score | Median human score |
|---|---|---|---|---|
| StealthGPT Super | 92% | 46 of 50 | 88.9% | 100% |
| WriteHuman | 65% | 32 of 49 | 63.6% | 91.39% |
| Rephrasy | 46% | 23 of 50 | 47.2% | 34.53% |
| Grubby | 36% | 18 of 50 | 39.8% | 10.69% |
| Undetectable.ai | 32% | 16 of 50 | 31.4% | 8.21% |
| RewriteAI | 28% | 14 of 50 | 29.6% | 0% |
| Ryne | 22% | 11 of 50 | 25.9% | 12.58% |
| HIX Bypass | 2% | 1 of 50 | 1.8% | 0% |
| Humbot | 0% | 0 of 50 | 1.4% | 0% |
| BypassGPT | 0% | 0 of 50 | 1.4% | 0% |
GPTZero 2026-09-13-base
| Humanizer | Strict bypass | Texts passed | Mean human score | Median human score |
|---|---|---|---|---|
| WriteHuman | 100% | 49 of 49 | 100.0% | 100% |
| StealthGPT Super | 98% | 49 of 50 | 99.0% | 100% |
| Undetectable.ai | 96% | 48 of 50 | 94.7% | 100% |
| Grubby | 94% | 47 of 50 | 93.6% | 100% |
| Rephrasy | 88% | 44 of 50 | 87.3% | 100% |
| RewriteAI | 86% | 43 of 50 | 83.3% | 100% |
| HIX Bypass | 68% | 34 of 50 | 63.6% | 94.97% |
| Humbot | 62% | 31 of 50 | 60.8% | 83.18% |
| BypassGPT | 56% | 28 of 50 | 57.4% | 67.06% |
| Ryne | 8% | 4 of 50 | 8.3% | 0% |
Originality.ai turbo
| Humanizer | Strict bypass | Texts passed | Mean human score | Median human score |
|---|---|---|---|---|
| Grubby | 98% | 49 of 50 | 98.2% | 100% |
| StealthGPT Super | 98% | 49 of 50 | 95.8% | 99.98% |
| WriteHuman | 98% | 48 of 49 | 96.1% | 99.98% |
| Undetectable.ai | 94% | 47 of 50 | 93.1% | 99.99% |
| RewriteAI | 88% | 44 of 50 | 91.3% | 99.99% |
| Rephrasy | 86% | 43 of 50 | 85.1% | 99.93% |
| Humbot | 70% | 35 of 50 | 69.1% | 72.15% |
| BypassGPT | 66% | 33 of 50 | 67.9% | 69.48% |
| HIX Bypass | 64% | 32 of 50 | 65.7% | 67.85% |
| Ryne | 0% | 0 of 50 | 2.1% | 0.13% |
ZeroGPT default
| Humanizer | Strict bypass | Texts passed | Mean human score | Median human score |
|---|---|---|---|---|
| BypassGPT | 100% | 50 of 50 | 97.2% | 100% |
| Humbot | 100% | 50 of 50 | 96.7% | 100% |
| HIX Bypass | 100% | 50 of 50 | 96.5% | 100% |
| Grubby | 100% | 50 of 50 | 96.0% | 100% |
| Undetectable.ai | 100% | 50 of 50 | 95.0% | 100% |
| WriteHuman | 100% | 49 of 49 | 94.8% | 98.2% |
| Rephrasy | 100% | 50 of 50 | 92.7% | 100% |
| StealthGPT Super | 98% | 49 of 50 | 91.1% | 97.2% |
| RewriteAI | 88% | 44 of 50 | 77.3% | 86.95% |
| Ryne | 62% | 31 of 50 | 56.6% | 59.65% |
Winston AI 5.0
| Humanizer | Strict bypass | Texts passed | Mean human score | Median human score |
|---|---|---|---|---|
| WriteHuman | 100% | 49 of 49 | 99.8% | 99.83% |
| Grubby | 100% | 50 of 50 | 99.3% | 99.85% |
| Rephrasy | 100% | 50 of 50 | 98.8% | 99.78% |
| BypassGPT | 100% | 50 of 50 | 98.0% | 99.63% |
| Humbot | 100% | 50 of 50 | 97.4% | 99.71% |
| StealthGPT Super | 98% | 49 of 50 | 97.7% | 99.81% |
| Undetectable.ai | 98% | 49 of 50 | 97.5% | 99.84% |
| HIX Bypass | 98% | 49 of 50 | 96.4% | 99.79% |
| RewriteAI | 96% | 48 of 50 | 95.4% | 99.69% |
| Ryne | 26% | 13 of 50 | 26.3% | 0.39% |
Quality
| Humanizer | |||||
|---|---|---|---|---|---|
| StealthGPT Super | 83.5% | 68% | 72% | 98% | 96% |
| WriteHuman | 58.7% | 22% | 49% | 98% | 65% |
| Grubby | 77.5% | 72% | 46% | 98% | 94% |
| Rephrasy | 54.5% | 36% | 30% | 92% | 60% |
| Undetectable.ai | 56.5% | 26% | 28% | 94% | 78% |
| RewriteAI | 72.0% | 54% | 40% | 100% | 94% |
| HIX Bypass | 46.5% | 18% | 6% | 96% | 66% |
| Humbot | 47.5% | 22% | 0% | 98% | 70% |
| BypassGPT | 49.0% | 22% | 2% | 94% | 78% |
| Ryne | 79.0% | 44% | 80% | 98% | 94% |
Downloads
scores.csv
50 rows of anonymized per-sample detector scores, output word counts and quality scores for all 10 tools.
summary.json
Aggregates, detector versions, settings used and quality pass rates.
Report README
The report as published in the data repo, under ID
2026-09-30-humanizer-comparison.
Data licensed CC BY 4.0; scoring code licensed MIT. Source and output text, internal IDs and raw API responses are withheld to limit adversarial reuse. That is a known gap, not a feature.
Verify
git clone https://github.com/StealthGPT-Labs/humanizer-benchmark
cd humanizer-benchmark
python3 scripts/verify_report.py reports/2026-09-30-humanizer-comparisonHow to cite
HumanizerEval, “September 2026 review”, published September 30, 2026, https://humanizereval.com/cycles/2026-09/.