Skip to content

Receipts Not Vibes AI  ·  Haverstraw, New York

2026-07-31

I Tried to Measure Whether AI Humanizers Work

68 repositories, two experiments — one that worked, one that failed — and thirteen years of my own writing as the test corpus.


Exhibit A — a receipt

AI Humanizer Skills · Field Survey

68 repositories measured · GitHub API · as measured 2026-07-30


Repositories measured68
Commits observed4,565
Stars, most-popular prose skill32,030
Stars, best-maintained prose skill2,683
Projects abandoned >90 days15
Stars pointing at abandoned code30,430
Published output evaluations found0

Correlation: stars vs. maintenanceρ = +0.197
Humanizer damage, detected blind83.3%
…with vendor's prescribed fix76.7%
Chance baseline50.0%
Popularity as a proxy for qualityworthless

Every figure above is produced by a script in a public repository, from live API data, and can be regenerated by anyone. n is stated for each result in the write-up, and it is small — read the confidence intervals, not just the rates. One of the two experiments failed; the post-mortem explaining why is published alongside the result that worked.

You may have come across a category of tool called an AI humanizer. You feed it AI-generated text, and it strips out a curated list of tells: the em-dashes, the "delve," the list of three, the opening that repeats your question back at you. On GitHub there are dozens of them, packaged as agent skills you can install in about a minute.

I went looking for the best one. I never found it. Instead, I found that nobody knows which one is best, including the people who build them — and that for a month I couldn't even tell which of my own tools was doing the work. This post is the record of that search: the measurements, the experiment that worked, and the one that failed.

The stars don't mean what you think they mean

I had an AI agent spin up a whole research experiment: it tested 68 of these projects against the GitHub API, measuring stars, forks, commit history, release tags, and days since the last push. Every figure in this post is as measured 2026-07-30, and the Python script that produced them is public, so you can rerun it and check both me and the agents.

First, the most popular prose skill has 32,030 stars. The one ranked second has 14,714 stars; yet, it hadn't been touched in 134 days. In third place was a repo with 14,166 stars, and it's been idle 191 days. Between them, those two dead projects hold 28,880 stars. Across the whole field, 15 of the 68 projects have sat untouched for more than 90 days, and represent 30,430 stars between them.

Not all 68 turned out to be humanizers, either. Fourteen were other kinds of tools that show up in the same searches and never touch a sentence. That leaves 54 actual prose skills: 24 for English, 20 for other languages, and 10 built for specific niches.

Across those 54, the correlation between star count and maintenance activity is statistically nothing. Statisticians measure this on a scale from −1 to +1, where +1 means the two things rise together and 0 means no connection at all. Two standard ways of computing it landed at +0.197 (Spearman) and −0.029 (Pearson). In plain terms: knowing a project's star count tells you almost nothing about whether anyone is still maintaining it.

This is hardly breaking news. After all, a star is one click by one logged-in account. The problem is that a star never expires and there's no downvote — and since these skills are free and install in a minute, a star doesn't even imply the person tried the tool, much less that they kept it. A star count tells you what got attention once. It tells you nothing about whether people are still using the tool.

Meanwhile, the best-maintained English prose skill in the field has a relatively paltry 2,683 stars; yet, this one was pushed the same day I had the agents measure. On the popularity list it sits way down; on every sign of life, it comes in first.

One disclosure before I go further, because this site is going to keep making a point of disclosures: in May 2026, before I even conceived of this experiment, I filed an issue on that best-maintained project. It was a compatibility report to the avoid-ai-writing repo by Conor Bronsdon, and Bronsdon repackaged the skill as an installable plugin in response. To be clear, a reporter on one issue is not the same as a contributor. Regardless, the survey ranks a project I have a small history with in first place, so you deserve to hear both facts in the same breath. And because the ranking is computed by a public script from public data, you can check for yourself that my involvement had no way to tip it.

The most popular repo rejected three reports of the same bug. Nobody measured it.

Every GitHub repository comes with a public issue tracker: a page where anyone with an account can report a problem, and the maintainer decides what to do with it. In the most popular humanizer's tracker, three separate people over six months filed the same report: when you run this tool on writing a human actually wrote, it damages it. One of the three reporters documented that his own human-written prose, after being "humanized," got flagged by a detector as 100% AI — the exact opposite of the advertised effect.

The maintainer closed all three reports as "not planned," and declined the last one on the grounds that the tool's voice-calibration feature already handles that case. In other words: you can hand the tool samples of your real writing, and the maintainer's position was that once it has those samples, it preserves your voice instead of damaging it.

That defense might even be true. The problem is that nobody has ever shown it. The agents went looking for a published test — any case of someone in this field running real human writing through their own tool and then measuring whether the result still sounded like the author. Across 68 projects and tens of thousands of stars, they found zero. Every one of these tools makes a promise about your voice, and not one of them has published evidence that it keeps that promise. Three users said the promise was broken; the maintainer said a feature covers it; and no measurement existed on either side of the argument.

So I ran one on myself.

The experiment that worked

I have thirteen years of my own writing — blog posts and school papers, 2012 through 2020. That corpus is the "human" side of everything below, and I'm telling you that up front rather than letting you discover it. The test is one author, and the author is me.

The agents pulled 16 paragraphs from that corpus by a rule fixed in advance, not by taste, because a cherry-picked sample proves nothing. They ran each paragraph through the leading humanizer twice: once cold, and once with samples of my writing supplied for calibration, which is exactly the fix the maintainer prescribed.

The judging was blind, and the judges were AI too: fresh model instances that had never seen the runs. Each judge got a pair, my real paragraph next to the machine-edited version, and had to pick which one the tool had touched. Every judging seat first answered control questions, simple checks that catch a careless judge, and any seat that failed got thrown out before anything was scored.

Chance is 50%. The judges that passed the controls picked out the machine-edited version 83.3% of the time cold (25 of 30 — the odds of that happening by luck are about 3 in 10,000) and 76.7% calibrated (23 of 30, about 1 in 200 by luck).

The three closed reports were right. When the tool runs on writing a human actually wrote, it damages the writing, and judges can see the damage. And the maintainer's defense — that calibration covers it — bought about six points. Mitigation, not coverage.

The mechanism is the part worth keeping. The tool barely touched the text: it changed a median of 8–9% of the words. Five to seven of the sixteen paragraphs came back more than 95% intact, and two came back completely untouched, byte for byte. The tool's own internal audit step kept reporting that the input already read human — and it edited anyway.

The judges caught it because of which eight percent the tool chose. My comma splices. An unclosed quotation mark. "In order to." The "that" in "I am convinced that." One judge wrote that a passage had "a casualness pass typical of humanizer tools." The tool treated the most identifying parts of my writing as the errors to fix. You can change eight percent of someone's prose and still be identified — by which eight percent you thought was broken.

The experiment that failed

I also built the opposite kind of tool. Instead of a list of things to remove, mine described what to reproduce: my recurring metaphors, how I structure an argument, the two registers I write in. Not a humanizer; closer to a portrait.

The agents tested four conditions: no tool, the portrait tool, the humanizer, and both chained together. In each condition, a mid-tier model drafted paragraphs on eight topics I'd genuinely written about, and the same kind of blind, control-checked judges tried to tell the drafts from my real writing.

The judges identified every arm, 16 out of 16, in all four conditions. And the arms did separate on one axis: how machine-like the text scored on a 0-to-10 tell scale. There, my portrait tool came in at 6.88 — worse than using no tool at all (5.38). My real paragraphs score 1.79 on the same scale.

The why is the useful part. The portrait described my peaks: the best metaphors, the sharpest structural moves. Given a document about peaks, the model wrote nothing but peaks. The real me opens a paper with "In this essay, I will evaluate two opposing scholarly articles." Then I hedge. Then I cite. Then I repeat myself. The judges' decisive cues were my plainness, my citation habits, and my mistakes.

I went back through the pattern lists of the projects the agents had surveyed. Not one of them flags too well written.

The failures taught more than a clean result would have

Three things went wrong with that second experiment, and I'd rather list them than have you find them.

The test couldn't rank anything. Every condition scored 16/16, including the no-tool baseline. When everything hits the maximum, the test can only tell you that nothing cleared the bar; yet, the bar was brutal: the drafts had to pass as my published prose, in a forced choice, in front of judges told to find the machine. My own rough drafts would fail that test against my finished ones. I built a test I could lose to myself.

I tested the wrong model. The drafting in the failed experiment ran on a mid-tier model, while everything that had actually impressed me in daily use had come from a frontier model. So the failed experiment answers a question nobody was asking. Of everything in the design, that choice mattered most, and I recorded it without registering what it meant until the numbers were in.

And the one that stung. When I had the agents audit which tools had actually run when, the audit showed that the month of writing I'd been proudest of was produced before the portrait tool existed. Zero invocations. The tool that had actually been running was the humanizer — the one I'd shrugged at and eventually switched off because I couldn't tell if it was doing anything. I keep full local logs of every session. I am about as instrumented as one person can be. I still credited the wrong tool for a month, because nothing in either tool's interface tells you what ran.

If you're picking one of these

  1. Sort by last commit, not stars. One click filters out two 14,000-star dead projects.
  2. Check that it's actually a skill, and that it touches prose. Some of the most-starred projects in these searches never see a sentence.
  3. Don't run a humanizer on writing you actually wrote. Three people reported the damage, the maintainer closed all three reports, and when I measured the damage on my own writing it was real at 83%.
  4. Distrust any "beats AI detectors" claim that comes without numbers. The projects that measured honestly report that surface rewriting barely moves modern detectors — and if you're a student, your school's AI-disclosure policy, not any tool, is the constraint that matters.
  5. If you build a voice tool, document the median sentence, not the highlights. Otherwise you get a caricature that reads more machine-made than the machine did.
  6. Before you A/B two writing tools, make sure you'll be able to tell which one ran. Attribution comes before quality. I learned that one the expensive way.

Why this is the first post on this site

This site is called Receipts Not Vibes AI. That was a rule scrawled at the top of my project notes for weeks before it was a domain name: decisions backed by evidence you can check, including, especially, when the evidence goes against you.

So this is the first receipt. One survey anyone can rerun. One result with its confidence intervals — the honest error bars — printed next to it. One failed experiment published with its post-mortem instead of buried. A list, in the repo, of the reasons my own results might be wrong. It's a small sample, on one author, and I've said so every time.

That's still more than a citation-free claim in a README. Everything is in the open: github.com/Daguilar0123/ai-writing-skill-field-guide.