Here’s the number that should worry anyone who writes for a living: between the first quarter of 2023 and the first quarter of 2026, Amazon’s self-published catalog grew 38.3 times over. Revenue grew 8.9 times.
That gap is the whole story. A far larger pile of books is now fighting over a money pool that barely budged in comparison.
A new analysis of 14,419 randomly selected self-published e-books, all released between January 2023 and March 2026, makes the case that AI-generated titles aren’t winning on quality. They’re winning on volume. And the collateral damage lands on books where no AI text was detected at all.
The data isn’t scraped guesswork
Most previous attempts to measure this problem sniffed at short book previews and guessed. This one didn’t.
The researchers classified every book on its full text using the Pangram v3.3 detector, whose developers report a false-positive rate of 0.04 percent. Pangram 4 has since shipped. Books landed in one of three buckets by share of text flagged as machine-written: none, light at up to 25 percent, and substantial above 25 percent.
Sales figures came from an internal dataset kept by one of the five major US publishers, tracking roughly 500,000 Amazon titles and covering about 95 percent of all e-books sold daily on the platform, according to the researchers.
The surface numbers look reassuring, and they’re misleading
Read the topline and you’d conclude AI books are flopping. Titles with substantial AI content are 20 percent of the catalog studied but pull just 12.1 percent of sales and 11.3 percent of revenue. Books with no detected AI text are 62.9 percent of the catalog and take 72.5 percent of revenue.
Case closed on the slop theory, right? Not quite.
The number of titles selling per quarter grew 19.2 times against that 8.9x revenue growth. Everyone’s slice got thinner, including the writers doing the work themselves.
Six of eight genres went backward
Comparing titles released in 2023 against those released in 2025 over the same post-release window, revenue per book fell in six of eight genres.
Narrow it to books with no detected AI text and it gets worse: revenue fell in seven of eight.
That’s the finding that kills the easy rebuttal. You can’t blame the drop on a flood of dud AI titles dragging down the average, because the human-written books are earning less on their own terms. The authors call the effect “dilution,” while stressing that their comparisons are observational and associational, not experimental proof of causation.
One genre bucked it. In Fantasy/Supernatural/Horror, where AI text showed up latest and stuck least, revenue per book for titles with no detected AI text climbed 35 percent. The researchers point to that reversal as an argument against blaming some broad market slump.
Kindle Unlimited genres show it sharpest
In genres with high Kindle Unlimited availability, where readers pull from one shared subscription pool, the revenue-share lead held by books with no detected AI text is 8.4 percentage points smaller than in low-availability genres.
The researchers credit genre-specific traits for the gap and explicitly decline to pin it on Kindle Unlimited itself.
The bestseller lists moved too. New Top 25 entries with substantial AI content went from near zero to 31 percent across the study period. Turnover at the top sped up: the share of books with no detected AI text holding a Top 25 spot from one quarter to the next fell to roughly 28 percent at one point before landing near 62 percent by the end.
A handful of accounts are doing most of the flooding
This isn’t thousands of hobbyists. Of 385 author identities that published more titles with substantial AI content after their first AI book, 287 raised their monthly output afterward.
The top-earning pseudonym pulled $1.7 million in gross revenue before platform fees across eight titles. The single highest-grossing book with substantial AI content made $643,000 on 80,431 copies sold.
It tracks with the New York Times report on “Coral Hart,” who reportedly put out more than 200 romance titles under 21 pen names in one year and sold about 50,000 copies. Books aren’t the only target either. One man scammed millions of dollars through streaming platforms using AI-generated songs.
The rare-phrase test is the uncomfortable part
To gauge how much the successful AI books overlap with language from existing works, the researchers ran the Allen Institute for AI’s infini-gram tool against the Google Books index.
They hunted for rare expressions appearing in five or fewer Google Books volumes and entirely absent from a 4.7-trillion-token web snapshot. Phrasing that specific points straight at published books rather than the open internet.
Among the 50 highest-grossing titles with substantial AI content, those rare expressions covered 45 percent of the text. For the top 50 books with no detected AI text, 37.7 percent. For award-winning or award-nominated fiction, 19.1 percent.
And within AI books, overlap climbed 7.6 percentage points for every tenfold jump in revenue. No such correlation turned up for books with no detected AI text. The method can’t trace where any individual passage came from or prove a specific book was copied. It measures aggregate language overlap, nothing more.
Why this matters in court
Chakrabarty told Ai2 that an AI detector only returns “an estimate—a score for how likely a passage is to be synthetic,” without pointing to where the language came from. When a suspect text also carries rare expressions absent from the web and present in only a handful of books, one “can say with some confidence that it was taken from books.”
That kind of evidence “acts as circumstantial evidence that supports an AI detector score” and simultaneously “helps debunk some hackneyed arguments that liken human reading of books to AI being trained on books,” a standard defense from AI companies in copyright fights.
In Kadrey v. Meta, Judge Vince Chhabria ruled for Meta in June 2025 but attached a pointed warning. He said it was hard to imagine that using copyrighted books to build a product generating billions in revenue while producing a potentially endless flood of competing works would qualify as fair use.
The plaintiffs there had no empirical evidence of market dilution to offer. This study is precisely that missing evidence: books with no detected AI text earning less as AI titles pour in.
Amazon knows and won’t tell you
Authors publishing through Kindle Direct Publishing must disclose AI involvement. Amazon doesn’t pass that disclosure on to customers.
So the one party with a clean list of which books are machine-written keeps it internal, while shoppers browse blind. The platform has wrestled for years with AI titles hijacking the names and styles of known authors, and its main answer has been a cap of three publications per day.
The findings on language overlap line up with a November 2025 study showing language models can reproduce passages from copyrighted books nearly word for word, and with another fall 2025 paper showing two books are enough to fine-tune a model on an author’s style.
If you write in a Kindle Unlimited-heavy genre and your per-title earnings have sagged since 2023 without your output or quality changing, you now have numbers behind the hunch. Take them to a lawyer, because Chhabria all but wrote the roadmap for what evidence a plaintiff needs.