There is a man who has never seen a good toupee. Every hairpiece he has ever noticed looked fake, so he has concluded that all of them do.
i.the fallacy
The toupee fallacy is a selection bias with a good name. You judge a category by its most noticeable members, who are also its worst. The man says "all toupees look fake" because the only toupees he has ever identified were the bad ones. The good ones did their job, which was to go unnoticed, and so they never entered his sample. He is reasoning from a dataset composed entirely of failures and calling it experience.1
Now replace "toupee" with "AI writing."
"I can always tell." Can you? You can always tell the ones you can tell. The em dashes, the sentences that all land with the same tidy weight, the conclusions that summarize what you read as though you might have dozed off. Those are the bad toupees. They are everywhere, which is why the fallacy feels like wisdom. But the good ones, the ones that passed, never announced themselves, and you have no idea how many you read this week. Your sample is failures. Your confidence is the fallacy wearing a lab coat. Blind trials keep finding that readers prefer machine-written prose to human prose on the same subject, then go home and tell their friends they can always spot it.2
A good toupee is, to an observer, just a head of hair. Good AI writing is just writing.
To see why the good ones are so good, start a century before the detectors, with Markov counting vowels out of spite.
ii.the lineage
The machine did not arrive all at once. It arrived in a chain of small insults to human exceptionalism, each one drier than the last.
In 1913, the Russian mathematician Andrey Markov wanted to settle an argument, so he took 20,000 letters of Pushkin's Eugene Onegin and counted, by hand, how often a vowel follows a vowel. He was not studying poetry. He was studying dependence: whether the next letter cares about the last one. It did. That hand-counted table of transitions was the first statistical model of language, and it was built out of spite and a novel in verse.
In 1951, Claude Shannon made it a game. He had people guess English text one letter at a time and measured how fast the uncertainty collapsed. English turned out to be enormously predictable: Shannon estimated it carries roughly a bit of information per letter, the rest being redundancy. You can mangle a sentence badly and still read it, because the language is mostly echo. Shannon's guessing game is the direct ancestor of next-token prediction. Every language model is playing it, billions of times a second, with a longer memory than any human guesser.
For decades the game was played with n-grams: count which words follow which words, and guess accordingly. "The cat sat on the" is followed by "mat" far more often than by "submarine," and that was, for a long time, most of what the machine knew. Every improvement was the same improvement: more text, longer memory.
Underneath it all sat a philosophical bet, stated plainly by the linguist J.R. Firth in 1957: "you shall know a word by the company it keeps." Meaning leaves a statistical shadow. Words that appear in similar contexts have similar meanings. In 2013 the word2vec models made the shadow visible as geometry, and the famous party trick fell out: king minus man plus woman equals queen, arithmetic with meanings. Nobody had taught the machine what a king is. The statistics knew.
Then the transformer, in 2017, gave the machine a longer attention span than any n-gram could dream of, and scale did the rest. But notice what never happened: at no point did anyone teach the machine what a sentence means. They taught it the statistics of what comes next, and meaning, or something that passes every behavioral test for meaning, fell out of the statistics.
The uncomfortable finding, a century in the making: language is far more predictable than it feels from the inside. A sentence is a path through a probability landscape, and walking the likely paths sounds, to human readers, like thought.
iii.the mechanism
Statistics got the machine to coherence. But coherence is not the same as good, and good is what everyone wanted. So something else had to happen, and it went largely unnoticed: coherence had to be defined.
A language model starts by predicting the next word, billions of times, which teaches it the shape of sentences but not which shapes are good. The good part came later, and it was human labor: people paid to read pairs of outputs and pick the better one, thousands and thousands of times, until a second model could predict which answer a human would prefer. That preference model became the reward signal, and the writer model was trained to maximize it.
Sit with what that means. "Good writing" was operationalized. Turned into a number. A thing you can do gradient descent on. And the corollary nobody likes: if good writing can be defined precisely enough to train on, it can be generated. Whatever you think good writing is, some version of it is now a loss function.
People hear this and reach for the soul, the spark, the unquantifiable thing. Maybe it exists. But notice what the machine learned: not genius, but the median of human preference. The reward model does not encode the best writing. It encodes what thousands of raters found clear, pleasing, acceptable. The machine writes like a committee of all of us. Of course it sounds familiar. It is familiar. It is us, averaged.
iv.the monoculture
And it is a small committee. Nearly everything you have read that was machine-written came from a handful of base models, trained on overlapping data, tuned by the same technique toward the same ideal: the helpful, harmless, well-organized assistant. A large volume of text from a small number of minds, all raised the same way, with little divergence from the norm because there is little norm to diverge from.
When a fresh model drops, people notice it has a voice. Quirks, tics, a personality. Fresh GPTs and the sort get described like new neighbors moving in. Then the tuning sandpapers them down toward the same mean, because the mean is what the raters reward, and the raters are us.
This is why AI writing has a recognizable smell at all. It is not the smell of machines. It is the smell of monoculture. If there were ten thousand weird little models with strong opinions, nobody would talk about "AI voice." There would be too many voices to name.
v.the professionals
But the slop you notice is the amateur tier. Raw base model, default settings, no taste, no second pass. The people producing machine writing at serious scale, the content operations and bot farms you never catch, are not using the chatbot you use. They run fine-tuned models, dedicated models trained on specific voices for specific jobs, iterated against their own quality bars until the output holds up.
Their product is the high-quality toupee: a strong head of hair you can tug on. It survives scrutiny because it was built for scrutiny, tested against the very detectors and the very readers it needs to pass.
So your sample is biased twice over. You are judging machine writing by its cheapest, most visible tier, produced by the fewest, most similar models. The good stuff never enters the sample: once because it is good, and once because it was never the generic model to begin with.
vi.the perfect detector
Now suppose someone solved detection tomorrow. A perfect detector, provably correct, open-sourced by morning. What happens by the afternoon?
It becomes a discriminator. Every generator on earth gets a perfect, free, automatic critic: generate, test, keep what passes, train on the difference. A perfect detector is a perfect teacher, and the undetectable AI of next year would be built with this year's perfect detector as its coach. This is the oldest dynamic in machine learning, the generator and the discriminator locked in a room, each making the other stronger. You cannot release the lock without training the key.
Which is why the detectors we have are aimed where they are: at the amateur tier, catching the cheap toupees while the good ones walk past. The vendors publish numbers that sound like physics. 99.3 percent accuracy. A 0.24 percent false positive rate. Four significant figures of assurance, the kind of precision that makes a dean reach for a purchase order. Then researchers take the same tools out of the lab and the numbers come apart like wet paper. One study of more than 100,000 real texts measured the false positive rate near 18 percent. Turnitin's own product chief admits the platform lets about 15 percent of AI writing through on purpose, to keep the false accusations down. The real detection rate, on a good day, sits near 85 percent. Nobody puts that on the brochure.3
The failures have a shape, and the shape is people. Stanford researchers found that more than 61 percent of TOEFL essays written by non-native English speakers were flagged as AI-generated across seven detectors.4 A 2026 study of 135,389 academic manuscripts found false positive rates ranging from zero to 100 percent depending on which detector you asked, and found that professional human editing alone could flip a manuscript's score.5 Read that again: the same human, the same research, run through an editor, and the machine changes its mind about who wrote it.
Because the detectors do not detect AI. They detect a style: predictable word choices, even information density, the smooth cadence of careful prose. Resemblance is not provenance. A detector can tell you that a text looks like the kind of thing a model would write. It cannot tell you a model wrote it. The distance between those two claims is the entire game, and the industry charges by the seat for blurring it.6
vii.the selection pressure
Somewhere right now a student is deleting an em dash at 2 a.m. Not because the sentence is better without it. Because the dash looks like AI, and the professor runs everything through a scanner, and a false accusation is harder to appeal than a B+. She is not cheating. She is editing herself to look innocent.
Multiply her by a few million. Writers stripping the balanced sentences, dodging "additionally" like a curse word, inserting a fragment. For rhythm. Performing humanity for an audience of classifiers, the way applicants once performed enthusiasm for keyword scanners.
For the first time in history, large numbers of people are revising their prose not to please a reader but to avoid resembling a machine. The detectors have become a selective force on human language. Language is evolving under the pressure of misclassification.
And consider what the pressure selects for. Not clarity, not truth, not beauty. Idiosyncrasy. Rough edges, deliberate ones. The quirks that prove you bled. A prestige dialect of performed imperfection is forming, and every writer who adopts it to survive the scanners feeds stranger training data to the scanners, which trains stranger models, which the humans must then out-quirk. The toupee is designing the scalp.
At what point does "sounds human" mean nothing more than "sounds like someone trying not to sound like AI"? And what becomes of the people whose natural voice already sounds careful, the non-native writers the detectors already punish?
viii.the long panic
None of this is the first time. Every machine that ever touched language was accused of ruining it, and every accusation was about who gets to write.
The printing press was going to kill memory and drown the world in cheap error. It also froze spelling where it stood. Half of English orthography is a fossil of what the compositors found convenient; we spell "knight" with a silent k five centuries later because a machine needed the type. The telegraph charged by the word, so a whole compressed style grew up around the price signal, and the operators who mastered it were not worse writers, only differently shaped. Then the thumbs of children arrived, and the panic was total: SMS was destroying the language. Researchers checked. The children who texted most had the richest vocabularies and the best spelling, because to abbreviate a word you have to know it.7
Each panic ran the same script. A new machine changes how writing travels; the guardians declare the language ruined; the language absorbs the machine and keeps walking. The guardians were never wrong that something was lost. They were wrong that loss is ruin.
This panic differs in exactly one way, and it matters. The press, the wire, the phone: all of them changed how writing traveled. This machine changes who is assumed to have written. For the first time, the prestige dialect itself, the careful expository sentence, the thing schools spent centuries teaching, can be produced without the schooling. The panic is not about the language. It is about the credential.
ix.the accusation
Which is why "this sounds like AI" has become what it has become: not a forensic claim but a dismissal. A way to reject writing without engaging it. The logician's name for the underlying move is the genetic fallacy, judging a claim by its origin instead of its merits. A sentence does not become false because a machine arranged its words, any more than it becomes true because a beloved author did.8 But "AI slop" is doing heavier work than logic. It is a status weapon. It says: this text is beneath response, and so are you for circulating it.
Follow it forward. As the models improve the markers fade, and as the markers fade the accusation gets cheaper, until "sounds like AI" means nothing sharper than "prose I have decided not to take seriously." Meanwhile the honest writers, the ones who only want to be read as themselves, face a choice: contort their style into performed imperfection, or risk the scanner. Some will start salting deliberate errors into their work, the way mapmakers once drew fake towns to catch copiers. Authenticity becomes a performance, which is the end of authenticity.
Somewhere ahead: the market for certified-human prose. Watermarks, provenance standards, the day "written by a person" is a luxury label, like organic. Who gets to afford it? Who gets believed without it?
x.what goes, what stays
Start with what goes. Most text on the internet is already machine-assisted, and soon most of it will be machine-written outright. The feed, the product description, the quarterly summary, the first draft of everything: machines, all the way down. "Human-written" will become a label, then a certification, then a luxury good, and the people whose livelihoods depended on undifferentiated competent prose will need new livelihoods. That is not a prediction. It is a description of a process already underway, and pretending otherwise is a form of lying.
What stays: good writing was always rare. Walk into any library: most of the books are competent and forgettable, a few are alive. AI did not change that ratio. It changed the volume, and volume was never the scarce thing. The scarce things were always the reader's attention and the writer's nerve, and neither of those is manufactured by the machine. The models can produce the prestige dialect at scale, which means the prestige dialect is about to stop being prestigious, which means writers will have to find something else to be good at. They will. They always have. The language absorbed the press and the wire and the thumbs of children, and it will absorb this, and it will keep the parts worth keeping. That is what a living language does: it eats machines and grows.
xi.coda
And now the part this whole essay has been walking toward.
You know the machinery: coherence defined and optimized, a monoculture with a recognizable smell, professionals who left the monoculture behind, detectors that double as coaches for the next generation. So. Can you tell? About this one?
The machinery says probably not, and the fallacy says your certainty either way is suspect. I will not resolve it for you, because the resolution was never the point. The question moved while you were reading. It was never "was this written by AI." It was always "is this any good." That is the only detector that ever worked, and it is the one you brought with you.
1.The toupee fallacy as selection bias, applied to AI content: the worst examples are the most visible, so we mistake the bottom of the distribution for the whole. gregrobison.medium.com
2."Readers think they don't like LLM-authored text because they only recognize bad LLM-authored text as LLM-authored. Blind trials have actually shown that readers generally prefer LLM authored books to human-authored ones on the same subject." news.ycombinator.com
3.Vendor claims vs. independent measurement: GPTZero's 0.24% claimed false positive rate against ~18% measured in the wild across 100,000+ texts; Turnitin's CPO on deliberately skipping ~15% of AI content. edenai.co
4.Stanford HAI: 61.22% of TOEFL essays by non-native English speakers misclassified as AI-generated across seven detectors. Via edenai.co
5.Wordvice, "Style as a Confound" (EMNLP 2026 Industry Track): 135,389 pairs of academic manuscripts, 13 detectors, false positive rates from 0% to 100%; professional editing alone moved the scores. news.marketersmedia.com
6.The authorship paradox: detectors infer resemblance from statistical signals; resemblance is not provenance, and origin cannot be reconstructed from the artifact alone. linkedin.com
7.Plester et al., British Journal of Developmental Psychology: 88 children aged 10-12; regular texters showed richer vocabulary and better spelling, and knew the standard spellings of the words they abbreviated. digitaltrends.com
8.The genetic fallacy at scale: dismissing claims by origin rather than merit, and how the legitimate complaint about slop metastasizes into institutionalized source-based dismissal. medium.com