In April, dozens of comments and posts going back 10 years vanished from r/AskHistorians. Not edited. Not hidden pending review. Removed, automatically, from a subreddit whose entire value proposition is that its answers stay readable for years.
“And there was nothing we or the experts [who posted the deleted content] could do about it,” said Dr. Sarah Gilbert, one of the community’s moderators.
That’s the version of AI moderation nobody puts in a press release. And it’s the one worth starting with, because the same month that happened, Reddit was telling everyone its AI had increased enforcement actions on hate and violent content by more than 200 percent.
What the mods actually saw
The Slack channel that receives links to modmail messages for the r/AskHistorians mod team is usually busy. That day it was flooded with alerts.
After recovering the text of some of the removed posts, one mod noticed a pattern: everything that got nuked linked to Rare Historical Photos, a historical image-sharing site. The team’s working theory is that Reddit flagged the domain as spam, and anything using its images as explanatory illustrations went down with it.
The mods believe Reddit’s recently revamped AI moderation tools were responsible. Reddit has not responded to a request for comment.
Here’s what makes it worse than a routine false positive. Gilbert told me some contributors spend hours, “sometimes over the course of days,” researching and writing a single answer. The subreddit’s users treat the place as an archive. Deleting a 10-year-old comment doesn’t inconvenience someone for an afternoon; it removes work that was still teaching people.
The numbers that don’t mean what they look like
Reddit says AI drives “faster, higher volume enforcement.” It said this month that AI has “helped reduce exposure to potentially harmful content by more than 40 percent.” It says its tools have “revoked nearly [2 million] fake votes daily,” and that it uses large language models “to catch the highly subtle, coordinated patterns of fake behavior and artificial hype that older systems once missed.”
Every one of those is a volume metric. None of them distinguishes a removed neo-Nazi screed from a decade-old citation to a photo archive. If the AskHistorians removals were logged as enforcement actions, they made the 200 percent figure look better.
Gilbert doesn’t trust the accounting, and she’s specific about why. “Back when there was more transparency in the system, we would routinely report hate and get an automated response that it wasn’t actually in violation of Reddit’s rules, prompting us to start an appeals process,” she said. “So it’s hard to trust the numbers because it’s hard to trust the ‘judgment’ of Reddit’s systems.”
False positives, she said, are a “huge problem” on Reddit.
Discord banned 8,400 people over chessboards
If you want the clearest illustration of how badly this fails without a person in the loop, Discord provided it. The company admitted its AI mod system wrongfully banned about 8,400 accounts from May to early July.
The cause: the AI classified images containing square grids as CSAM. Chessboards. Spreadsheets. Upload one, get a permanent ban. Discord says all affected accounts have since been reinstated.
The company’s explanation is the part I keep coming back to. Discord said its AI moderation was never meant to run unsupervised, and that a human employee is supposed to review AI-flagged content before any action is taken. A bug let the AI skip the human step and ban people directly.
So the safeguard existed on paper. It took a software defect and roughly two months for 8,400 accounts to find out it wasn’t running.
The problem AI created, and is now sold as the fix for
Generative AI didn’t just break moderation from the enforcement side. It flooded the input side first.
Large language models “have made spam detection a lot harder,” Gilbert said, because they’re built to imitate real human voices. “Over the last two to three months, we’ve been absolutely flooded by LLM-powered spambots,” she said.
There’s a commercial engine behind some of that. Marketing agencies now produce social posts designed to get brands cited by chatbots. Inauthentic posting for visibility is old; aiming it at chatbot outputs is new. The startup ReachLLM focuses specifically on marketing through chatbots, and its representatives have created and moderate subreddits on Reddit.
Who gets caught in the net
Typical moderation systems use machine learning classifiers to scan posts and flag rule-breaking content. Sarcasm, satire and slang are exactly the things a classifier handles worst.
That’s not a neutral failure. Research suggests marginalized groups are disproportionately affected by AI moderation. Gilbert, who is also research director of Cornell’s Citizens and Technology Lab, said “marginalized and vulnerable populations are among those who experience the highest rates of moderation, and that typically this is a result of ‘false-positives,'” frequently triggered by counter-speech, language reclamation and “responses to hateful content.”
“False positives are an equity issue. They mean that groups that are already marginalized are further silenced and censored,” she said.
Read that against the 200 percent enforcement increase. The systems built to protect vulnerable communities from hate are penalizing those communities for talking back to it.
Meta and Tumblr have the same bug
Since 2025, Facebook and Instagram users have complained about mass bans they attribute to AI moderation. Meta hasn’t said whether AI is behind them. What it has done is lean harder on generative-AI-based moderation instead of human reviewers, a shift some people, including Meta employees, say is moving too fast.
The absence of a human on the other end is its own injury. Users who say they broke no rules have had no way to reach a Meta employee to find out what happened or how to get reinstated.
Tumblr has its own version. In March, Chenda Ngak, head of communications at Tumblr parent company Automattic, said Tumblr’s automated systems wrongfully banned “sub-200” accounts in a single afternoon. In 2025, users complained that automatic moderation misflagged content as “mature,” cutting its reach. Tumblr never confirmed AI caused either problem, though it has said it uses “a mix of machine-learning classification and human moderation.”
The quieter cost: mods lose the call
There’s a side effect that doesn’t show up in any transparency report. Some subreddit moderators would rather ban a user outright for hateful or violent rhetoric. But when Reddit’s AI removes the content before a human mod ever sees it, the mod can’t judge whether a ban is warranted.
Automated removal looks like a win in the logs. On the ground, it strips the community of the context it needs to police itself.
Reddit is at least moving on this. This week it expanded testing for Rules Hub, which lets human mods “choose which rules should be automatically enforced, decide what happens when a rule is triggered (send to queue, filter, or remove), preview the experience before enabling it, and review logs and insights.” Reddit expects it to eventually replace Automod, which leans on exact keyword matching.
That’s the right direction: more control for the people who know the community, not less.
What would actually fix this
AI moderation saves platforms money and pulls harmful content down faster than humans can. Both true. But a system that can’t tell a checkerboard from CSAM has not earned the right to act alone, and Discord’s own account of what went wrong is the argument for keeping a person in the loop.
Mods I’ve spoken with keep pointing at the generative AI boom as the reason rule-breaking content is spiking. For platforms that run on what users contribute, that’s an existential problem, not a support ticket.
The answer isn’t less AI. It’s machine-scale detection paired with human judgment that has the authority to override it, and the transparency for mods to check the machine’s work. Reducing human input while the volume of low-effort AI content climbs is moving backward on purpose.
Gilbert’s team spent April recovering text from posts that took contributors days to write. Nobody at Reddit had to sign off on deleting them. That’s the design flaw.
Advance Publications is the largest shareholder in Reddit.