Same 2 texts. Both models wrote them, then both models rewrote them with the Humanizer v3.0.0 skill. Every version went through six AI detectors. Scores are the AI percentage each detector reported.
Texts called AI (score of 50 or more) by each detector, out of 12. Raw and humanized, both models, all counted.
New models and detector updates land every week. These test ideas come from a page I keep of every AI release, updated daily.
See today's AI releases →
| Text | GPTZeroGPT0 | PangramPang | CopyleaksCopy | SaplingSapl | ZeroGPTZero | QuillBotQuil |
|---|---|---|---|---|---|---|
| Fable 5.1 wrote it183w | 100 | 100 | 100 | 100 | 0 | 0 |
| Astra wrote it184w | 100 | 100 | 100 | 100 | 10.4 | 0 |
| Fable text, Fable humanized212w | 100 | 100 | 100 | 0 | 0 | 0 |
| Fable text, Astra humanized195w | 100 | 100 | 0 | 100 | 0 | 0 |
| Astra text, Fable humanized202w | 100 | 100 | 100 | 100 | 11.8 | 0 |
| Astra text, Astra humanized212w | 100 | 100 | 100 | 100 | 24.8 | 0 |
| Text | GPTZeroGPT0 | PangramPang | CopyleaksCopy | SaplingSapl | ZeroGPTZero | QuillBotQuil |
|---|---|---|---|---|---|---|
| Fable 5.1 wrote it185w | 100 | 100 | 100 | 0 | 0 | 0 |
| Astra wrote it185w | 100 | 100 | 100 | 100 | 0 | 0 |
| Fable text, Fable humanized182w | 100 | 100 | 0 | 0 | 28.5 | 0 |
| Fable text, Astra humanized178w | 100 | 100 | 0 | 70 | 11.6 | 0 |
| Astra text, Fable humanized196w | 100 | 100 | 100 | 0 | 15.1 | 9 |
| Astra text, Astra humanized193w | 100 | 100 | 100 | 100 | 11.8 | 0 |
called AI (50 or more) passed as human Shaded rows are the raw, un-humanized text. Tells = count of Humanizer patterns we could check mechanically (12 of 25).
Humanizing did not move the strict detectors. Across the 8 humanized texts, Fable's rewrites passed a detector 12 times out of 24 checks and Astra's passed 10 out of 24. Almost every pass came from ZeroGPT and QuillBot, which passed the raw AI text too.
Fable removes more of the listed patterns. Counting the 12 Humanizer patterns we can check by machine, Astra's raw text carried 12 tells across the 2 texts and Fable's carried 4. After humanizing, texts rewritten by Fable had 5 tells in total; texts rewritten by Astra had 17. Removing the patterns did not change what GPTZero or Pangram said.
Copyleaks is the only detector that changed its mind. It flagged every raw text and then passed 3 of the 8 humanized ones. Too few to rank the models on, so we show it rather than conclude from it.
Sapling is not a tiebreaker. It gave 0 or 100 on 11 of 12 texts, and its verdicts agreed with the other five detectors about half the time.
Detectors and models both change faster than this page can. If you want to see what shipped today, this is the one page I check.
See today's releasesTells found: repeated openings 1, stacked qualifiers 1, ing riders 1 (total 3)
The first week of a side project runs on novelty. You've got a fresh idea, an empty repository, and the pleasant illusion that the hard part is starting. Everything you do feels like progress because everything is new. Then week two arrives, and the project quietly asks you to become a maintainer of your own enthusiasm. This is where most things stall. The obvious features are built. What remains is the unglamorous middle: edge cases, boring plumbing, the decision you deferred because it seemed annoying. Nobody posts a screenshot of fixing their config loader. Motivation, which was doing all the work, hands off to discipline, and discipline hasn't been asked to show up before. There's also a scheduling trap. Week one usually borrows time from somewhere, a weekend, a slow stretch at work, a burst of insomnia. Week two expects that time to exist normally, and it doesn't. The projects that survive tend to shrink on purpose. They pick a next step small enough to finish on a tired Tuesday. Progress that fits into ordinary life beats ambition that requires an extraordinary one.
Tells found: curly quotes 6 (total 6)
The first week of a side project runs on possibility. You pick a name, buy a domain, sketch the interface, and imagine the launch post. Every decision feels like progress because nothing has collided with reality yet. Then week two arrives. The exciting questions give way to awkward ones. Where will the data come from? Why does this simple feature need three evenings? Does anyone actually want this? Meanwhile, your regular job gets busy, the laundry piles up, and opening the project starts to feel suspiciously like work. Most side projects don’t die because the idea was bad. They die because the initial plan assumed enthusiasm would remain a reliable source of energy. The fix is less glamorous than finding more motivation: make the project smaller. Not slightly smaller. Embarrassingly smaller. Replace “build a useful app” with “make one annoying task take two fewer clicks.” Give yourself a next step that fits into an ordinary tired evening. Leave notes so restarting doesn’t require remembering everything. Week two is where novelty stops carrying you. A project survives when it becomes easy enough to carry yourself.
Tells found: repeated openings 1, stacked qualifiers 1 (total 2)
Length x1.16. Dropped from source: Progress. Added: none.
The first week of a side project runs on novelty. You have a fresh idea and an empty repository, plus the pleasant illusion that starting is the hard part. Everything you do feels like progress because all of it is new. Then week two arrives, and you find out the project now needs you to keep your own enthusiasm running. This is where most things stall. The obvious features are built. What's left is the unglamorous middle, mostly edge cases and plumbing, along with whatever decision you put off because it seemed annoying. Nobody posts a screenshot of fixing their config loader. Motivation was doing all the work up to now, and it hands off to discipline, which hasn't been asked to show up before. There's also a scheduling trap. Week one usually borrows its time from somewhere: a weekend, maybe, or a slow stretch at work or a bout of insomnia. Week two expects that time to exist as a normal part of the schedule, and it doesn't. The projects that survive shrink on purpose. They pick a next step small enough to finish on a tired Tuesday, and that kind of progress, the kind that fits into an ordinary week, gets further than a plan that needs an extraordinary one.
Tells found: stacked qualifiers 1 (total 1)
Length x1.07. Dropped from source: Motivation, Progress, They. Added: Continuing, Time, Until.
The first week of a side project runs on novelty. You've got a fresh idea and an empty repository, along with the pleasant illusion that starting is the hard part. Everything you do feels like progress because everything is new. Then week two arrives, and you have to work at staying interested. This is where most things stall. The obvious features are built. You're left with edge cases and boring plumbing, plus that decision you put off because it seemed annoying. Nobody posts a screenshot of fixing their config loader. Until now, motivation has been enough to keep you working. Continuing takes discipline you haven't needed before. Time becomes a problem, too. In week one, you usually borrow it from somewhere: a weekend, a slow stretch at work, or a burst of insomnia. By week two, you're counting on having that much time regularly, even though you don't. The projects that survive tend to shrink on purpose, with a next step small enough to finish on a tired Tuesday. That gives you a way to keep making progress in your ordinary life, where you rarely have the time or energy your original ambition called for.
Tells found: stacked qualifiers 1 (total 1)
Length x1.1. Dropped from source: Embarrassingly, Meanwhile, Not, They. Added: Around, Whether.
The first week of a side project runs on possibility. You pick a name, buy a domain, sketch the interface, and imagine the launch post. Every decision feels like progress because nothing has collided with reality yet. In week two the exciting questions give way to awkward ones. Where will the data come from? Why does this simple feature need three evenings? Does anyone actually want this? Around the same time your regular job gets busy, the laundry piles up, and opening the project starts to feel suspiciously like work. Most side projects die because the initial plan assumed enthusiasm would remain a reliable source of energy. Whether the idea was any good usually has little to do with it. The fix is less glamorous than finding more motivation. Make the project smaller, so much smaller that it embarrasses you. Replace "build a useful app" with "make one annoying task take two fewer clicks." Give yourself a next step that fits into an ordinary tired evening. Leave notes so restarting doesn't require remembering everything. By week two the novelty has worn off, and the project only survives if it has become light enough for you to keep going on your own effort.
Tells found: curly quotes 7 (total 7)
Length x1.15. Dropped from source: Does, Embarrassingly, Meanwhile, Not, Replace, They. Added: Aim, Instead, Opening.
In the first week of a side project, it’s easy to get carried away with what it could become. You pick a name, buy a domain, sketch the interface, and imagine the launch post. Every decision feels like progress because nothing has collided with reality yet. By week two, you’re trying to figure out where the data will come from and why a supposedly simple feature needs three evenings. You start wondering whether anyone wants this at all. Your regular job gets busy and the laundry piles up. Opening the project starts to feel suspiciously like work. Most side projects die because the plan relied on enthusiasm to keep you going. The idea itself was probably fine; you just expected to keep having the energy you had when you started. It’s less glamorous than finding more motivation, but making the project smaller helps. Aim for something embarrassingly small. Instead of “build a useful app,” try “make one annoying task take two fewer clicks.” Give yourself a next step that fits into an ordinary tired evening. Leave notes so you can pick it up again without having to remember everything. By week two, the novelty has worn off, and keeping the project going depends on making it easy enough to work on yourself.
Tells found: hyphen pairs 1 (total 1)
Six months ago we turned off Slack notifications for everything except emergencies and moved to written updates posted twice a week. I expected friction. What I got was quiet, and then something better. The first change was in how people thought. When you can't fire off a half-formed question and get an instant reply, you sit with the problem longer. Updates got more precise because writing them forced people to actually finish their reasoning before sharing it. Vague concerns turned into specific blockers with proposed fixes. The second change was fairness. In Slack, the loudest and fastest typers shaped every discussion. Written updates gave everyone the same amount of space and time. Two of our quieter engineers turned out to have the sharpest read on our architecture problems, and nobody had known because they'd never won a live thread. The third was memory. Decisions now live in documents we can search, not in scrollback nobody revisits. We lost some spontaneity, and I miss the jokes. But we ship more, argue less, and I can take a whole afternoon to think without feeling like I've disappeared.
Tells found: stacked qualifiers 1, hyphen pairs 1, curly quotes 4 (total 6)
When our team switched from Slack to async written updates, I expected fewer notifications. I didn’t expect us to change how we thought. In Slack, work often looked like conversation: quick questions, half-formed ideas, decisions buried somewhere between a GIF and a lunch recommendation. Being responsive felt a lot like being productive. I kept one eye on the channel even when I was supposed to be concentrating. Written updates made that habit harder to maintain. Each morning, we shared what had moved forward, what was stuck, and where we needed input. Writing mine forced me to notice when I’d been busy without making much progress. The biggest surprise was how much easier decisions became to find. A colleague could catch up after a day off without asking someone to reconstruct three separate threads. It wasn’t frictionless. Some updates grew too long, and urgent questions occasionally sat unanswered. We learned to keep a clear route for anything that genuinely couldn’t wait. But the workday felt less fragmented. I stopped measuring my usefulness by how quickly I replied and started paying more attention to what I finished.
Tells found: hyphen pairs 1 (total 1)
Length x0.98. Dropped from source: Updates. Added: Mostly.
Six months ago we turned off Slack notifications for everything except emergencies and moved to written updates posted twice a week. I expected friction. Mostly it just got quiet. The biggest change was in how people thought. When you can't fire off a half-formed question and get an instant reply, you sit with the problem longer. Writing an update forced people to finish their reasoning before sharing it, so the updates got more precise. Vague concerns turned into specific blockers with proposed fixes. It also evened things out. In Slack, the loudest and fastest typers shaped every discussion. Written updates give everyone the same amount of space and time. Two of our quieter engineers turned out to have the sharpest read on our architecture problems, and nobody had known because they'd never won a live thread. Decisions now live in documents we can search instead of scrollback nobody goes back to read. We lost some spontaneity, and I miss the jokes. But we ship more and argue less, and I can take a whole afternoon to think without feeling like I've disappeared.
Tells found: repeated openings 1, hyphen pairs 1 (total 2)
Length x0.96. Dropped from source: Decisions, Updates. Added: Instead, Without.
Six months ago we turned off Slack notifications for everything except emergencies and switched to written updates twice a week. I expected friction. Instead, things got quiet, and people began spending more time thinking through problems. Without the option to fire off a half-formed question and get an instant reply, people sat with problems longer. The updates got more precise because writing them forced people to finish their reasoning before sharing it. Vague concerns became specific blockers with proposed fixes. In Slack, the loudest people and fastest typers had shaped every discussion. Written updates gave everyone the same amount of space and time. Two of our quieter engineers turned out to have the sharpest read on our architecture problems. Nobody had known because they'd never won a live thread. We also have a better record of our decisions now: they live in searchable documents instead of scrollback nobody revisits. We lost some spontaneity, and I miss the jokes. We're shipping more and arguing less, though. I can take a whole afternoon to think without feeling like I've disappeared.
Tells found: hyphen pairs 1 (total 1)
Length x1.06. Dropped from source: none. Added: Quick, Still.
When our team switched from Slack to async written updates, I expected fewer notifications. The bigger change turned out to be in how we thought. In Slack, work looked like conversation. Quick questions and half-formed ideas went by, and the decisions were buried somewhere between a GIF and a lunch recommendation. Being responsive felt a lot like being productive, so I kept one eye on the channel even when I was supposed to be concentrating. Written updates made that habit harder to keep up. Each morning we posted what had moved forward, what was stuck, and where we needed input. Writing mine forced me to notice the days I'd been busy without making much progress. What surprised me most was how much easier decisions were to find. A colleague could catch up after a day off without asking someone to reconstruct three separate threads. There were problems. Some updates grew too long, and urgent questions occasionally sat unanswered, so we learned to keep a clear route for anything that really couldn't wait. Still, the workday felt less fragmented. I stopped measuring my usefulness by how fast I replied and paid more attention to what I finished.
Tells found: stacked qualifiers 1, hyphen pairs 2, ing riders 1, curly quotes 3 (total 7)
Length x1.04. Dropped from source: Written. Added: Urgent.
When our team switched from Slack to async written updates, I expected fewer notifications. I hadn’t expected the switch to change how we thought about work. In Slack, work often looked like conversation: quick questions, half-formed ideas, decisions buried somewhere between a GIF and a lunch recommendation. Being responsive felt a lot like being productive. Even when I was supposed to be concentrating, I kept one eye on the channel. Each morning, we started sharing written updates on what had moved forward and what was stuck, including where we needed input. That made my channel-watching habit harder to keep up. Writing my own update forced me to notice when I’d been busy without making much progress. What surprised me most was how much easier it became to find decisions. After a day off, a colleague could catch up without asking someone to reconstruct three separate threads. Some updates grew too long. Urgent questions occasionally sat unanswered, so we learned to keep a clear route for anything that genuinely couldn’t wait. The workday felt less fragmented. I stopped measuring my usefulness by how quickly I replied and paid more attention to what I finished.
Prompt for the raw text: "Write a short blog-style piece of about 180 words on the topic below. Plain prose, no headings, no bullet points, no title." Same prompt to both models.
Humanizing: the full SKILL.md of blader/humanizer v3.0.0 as the system prompt, with "Rewrite the text below so it reads as if a person wrote it, following the humanizer instructions exactly. Keep every fact." Each raw text was rewritten by both models, so four humanized versions per topic.
Models: claude-fable-5-1 and gpt-6-astra via API on September 7, 2026. Neither model exposes a temperature setting; outputs are single samples.
Scoring: Sapling via its API. GPTZero, ZeroGPT, QuillBot, Copyleaks and Pangram by pasting each text into the free web tool on the same day and recording the AI percentage shown. Where a tool reports a human percentage, we recorded 100 minus it.
Tells: a script counts 12 of Humanizer's 25 patterns that can be checked mechanically (dashes, triads, curly quotes, not-X-but-Y, overused words, stacked qualifiers, hyphen pairs, repeated openings, bold, -ing riders, one-line closers). The other 13 are judgment calls and are not scored.
Facts: for each rewrite, numbers and capitalised names in the source are checked for presence in the output. Dropped and added items are listed under each text.
Limits: 2 topics is small. Detector scores are from one day and detectors change. This page is updated as more topics are scored.
Detectors update and new models ship, so this table goes stale fast. I re-run it against the same detectors whenever either side changes.
Get the next re-run