September 2026 review · same 50 texts

StealthGPT Super vs Grubby

In the September 2026 HumanizerEval review, StealthGPT Super and Grubby rewrote the same 50 English texts. StealthGPT Super cleared all 5 detectors on 46 of them; Grubby on 18. The difference is mostly Pangram: on 29 samples StealthGPT passed it where Grubby failed, and the reverse happened on 1. On GPTZero, Originality.ai, ZeroGPT and Winston AI the two were close to even.

Per-sample agreement

Per-sample agreement, September 2026. Counts are out of the 50 shared samples, computed at the 50% human-score threshold. Margin counts the samples only one tool passed, higher count first.
DetectorPassed more oftenMarginBoth passedNeither passed
Pangram 4.0StealthGPT29 to 1173
GPTZero 2026-09-13-baseStealthGPT3 to 1460
Originality.ai turboTied1 to 1480
ZeroGPT defaultGrubby1 to 0490
Winston AI 5.0Grubby1 to 0490
All 5 detectorsStealthGPT29 to 1173

Every count is computed from scores.csv at build time, not transcribed. “Neither passed” counts the texts that defeated both tools; on Pangram that was 3 of 50 here.

Quality and output discipline

Quality and output discipline, September 2026.
MeasureStealthGPT SuperGrubby
68%72%
72%46%
98%98%
96%94%
1.041.11
00
00
Super, production outputAcademic
Websitestealthgpt.ai grubby.ai

On per-sample quality scores, StealthGPT scored higher on 17 samples, Grubby on 9, and the two tied on 24. StealthGPT’s output was shorter on 35 of 50 samples; mean length ratios were 1.04 and 1.11, or 4% longer than the input and 11% longer than the input respectively.

Where StealthGPT outperforms Grubby

StealthGPT

  • Pangram: passed 46 of 50 samples, against 18 for Grubby.
  • GPTZero: passed 49 of 50 samples, against 47 for Grubby.
  • Naturalness: read fluently in 72% of outputs, against 46% for Grubby.
  • Stance preservation: preserved the source’s stance in 96% of outputs, against 94% for Grubby.
  • Length: shorter than Grubby on 35 of 50 samples, which matters under a word limit.

Grubby

  • Pangram: failed 32 of 50 samples, 29 of which StealthGPT passed.
  • GPTZero: failed 3 of 50 samples, 3 of which StealthGPT passed.
  • Naturalness: had awkward or clunky phrasing in 54% of outputs.
  • Stance preservation: distorted the source’s stance in 6% of outputs.
  • Length: longer than StealthGPT on 35 of 50 samples, so more to cut under a word limit.

StealthGPT vs Grubby FAQ

Is StealthGPT better than Grubby?

Yes. In the September 2026 HumanizerEval review, StealthGPT Super cleared all 5 detectors on 46 of the 50 shared samples and Grubby on 18. On GPTZero, Originality.ai, ZeroGPT and Winston AI the two were within 6 points of each other. Quality ran 17 samples to 9, with 24 ties.

Which humanizer is best for Pangram, StealthGPT or Grubby?

StealthGPT Super. In September 2026, StealthGPT Super passed Pangram on 46 of the 50 shared samples and Grubby on 18. On 29 samples StealthGPT passed where Grubby failed; the reverse happened on 1. Both failed on 3.

Which humanizer changes the text less?

StealthGPT Super. It scored higher on per-sample quality on 17 of 50 samples, against 9. StealthGPT Super’s outputs averaged 1.04× the input length and passed factual consistency on 68%; Grubby’s averaged 1.11× and passed on 72%. On stance preservation the split was 96% to 94%. StealthGPT produced the shorter output on 35 of 50 samples.

More comparisons