September 2026 review · same 50 texts

Grubby vs Rephrasy

In the September 2026 HumanizerEval review, Grubby and Rephrasy rewrote the same 50 English texts. Grubby cleared all 5 detectors on 18 of them; Rephrasy on 20. The difference is mostly Originality.ai: on 7 samples Grubby passed it where Rephrasy failed, and the reverse happened on 1. On GPTZero, ZeroGPT and Winston AI the two were close to even.

Per-sample agreement

Per-sample agreement, September 2026. Counts are out of the 50 shared samples, computed at the 50% human-score threshold. Margin counts the samples only one tool passed, higher count first.
DetectorPassed more oftenMarginBoth passedNeither passed
Pangram 4.0Rephrasy9 to 41423
GPTZero 2026-09-13-baseGrubby3 to 0443
Originality.ai turboGrubby7 to 1420
ZeroGPT defaultTied0 to 0500
Winston AI 5.0Tied0 to 0500
All 5 detectorsRephrasy9 to 71123

Every count is computed from scores.csv at build time, not transcribed. “Neither passed” counts the texts that defeated both tools; on Pangram that was 23 of 50 here.

Quality and output discipline

Quality and output discipline, September 2026.
MeasureGrubbyRephrasy
72%36%
46%30%
98%92%
94%60%
1.111.13
03
00
Academicv4
Websitegrubby.ai rephrasy.ai

On per-sample quality scores, Grubby scored higher on 32 samples, Rephrasy on 6, and the two tied on 12. Rephrasy’s output was shorter on 28 of 50 samples; mean length ratios were 1.11 and 1.13, or 11% longer than the input and 13% longer than the input respectively.

Where each one outperforms the other

Grubby

  • Pangram: passed 4 samples where Rephrasy failed, even though it was detected more often overall, passing 18 to 23.
  • GPTZero: passed 47 of 50 against 44, including 3 samples where Rephrasy failed, against 0 the other way.
  • Originality.ai: passed 49 of 50 against 43, including 7 samples where Rephrasy failed, against 1 the other way.
  • ZeroGPT: tied at 50 of 50 but on a higher mean human score, 96.0% to 92.7%. It passed 0 samples where Rephrasy failed.
  • Winston AI: tied at 50 of 50 but on a higher mean human score, 99.3% to 98.8%. It passed 0 samples where Rephrasy failed.
  • Factual consistency: 72% against 36%.
  • Naturalness: 46% against 30%.
  • Syntax integrity: 98% against 92%.
  • Stance preservation: 94% against 60%.

Rephrasy

  • Pangram: passed 23 of 50 against 18, including 9 samples where Grubby failed, against 4 the other way.
  • Originality.ai: passed 1 sample where Grubby failed, even though it was detected more often overall, passing 43 to 49.
  • Length: shorter output on 28 of 50 samples, which matters under a word limit.

Grubby vs Rephrasy FAQ

Is Grubby better than Rephrasy?

In the September 2026 HumanizerEval review, Grubby cleared all 5 detectors on 18 of the 50 shared samples and Rephrasy on 20. On GPTZero, ZeroGPT and Winston AI the two were within 6 points of each other. Quality ran 32 samples to 6, with 12 ties.

Which humanizer is best for Pangram, Grubby or Rephrasy?

Rephrasy. In September 2026, Grubby passed Pangram on 18 of the 50 shared samples and Rephrasy on 23. On 4 samples Grubby passed where Rephrasy failed; the reverse happened on 9. Both failed on 23.

Which humanizer changes the text less?

Grubby. It scored higher on per-sample quality on 32 of 50 samples, against 6. Grubby’s outputs averaged 1.11× the input length and passed factual consistency on 72%; Rephrasy’s averaged 1.13× and passed on 36%. On stance preservation the split was 94% to 60%. Rephrasy produced the shorter output on 28 of 50 samples.

More comparisons