Four sources described the same near-miss to CNN: a US Special Operations Command analyst fed intelligence reports about a Chinese ship’s manifest into a chatbot, and the chatbot got the cargo wrong. Not slightly wrong. The resulting report claimed the vessel was moving nuclear arms program components through the Middle East.
The military was preparing to intercept and board it, with air support ready, when officials worked out that the chatbot had “inaccurately identified the material the ship was carrying.” The intelligence was, per CNN’s report, “entirely false.”
One source put it more bluntly. The episode “almost started a war.”
What the chatbot actually did
This wasn’t a case of an analyst asking a general-purpose assistant to summarize a PDF. The chatbot “fused together open-source intelligence with secret signals intelligence in government holdings,” and that fusion went into a finished intelligence product.
That detail matters more than the headline does. Mixing open-source material with classified signals intelligence is exactly the kind of correlation work that looks like a natural fit for a language model and is exactly where a model with thin context will invent the connective tissue. It doesn’t flag the gap. It fills it.
We’ve watched this failure mode for three years
“Hallucinating” was the Cambridge Dictionary’s word of the year in 2023. Since then the list of professionals caught out by confident fabrication has gotten long and unflattering: non-fiction authors, journalists, academic researchers, judges, doctors, police departments, corporate call centers.
The usual mitigation is a prompt politely asking the model not to make things up. Some researchers suggest it may be impossible to prevent LLMs from hallucinating altogether.
So the question isn’t whether the Pentagon knew this could happen. It’s what the Pentagon did with that knowledge.
The answer is: it accelerated
In January the Department of Defense rolled out an “AI acceleration strategy” aimed at making “all appropriate data available across federated IT systems for AI exploitation, including mission systems across every service and component.”
“AI is only as good as the data that it receives, and we’re going to make sure that it’s there,” Defense Secretary Pete Hegseth said in rolling out the initiative.
Read that line against the ship incident. The data wasn’t the problem. The manifest existed. The model misread what was on it.
Three vendors, 1.5 million users
Last December the department announced it would build its “GenAI.mil” platform on Google’s Gemini for Government. Last month it added Grok for Government as an option. Anthropic offers a customized version of Claude for US spy work.
In June a Pentagon representative told Congress, with some pride, that the department uses generative AI to help produce congressionally mandated reports, and that 1.5 million active DoD personnel have used the military’s generative AI tools.
That’s the scale worth sitting with. One analyst’s chatbot query nearly put a boarding party on a Chinese vessel. The same tooling is in front of a million and a half people.
The safety language hasn’t aged well
A 2023 State Department “Declaration on Responsible Military Use of Artificial Intelligence and Autonomy” said “principled” military use of AI “should include careful consideration of risks and benefits, and it should also minimize unintended bias and accidents.” It insisted that “accountable” use must always involve “a human in the loop, a responsible human chain of command and control.”
There were humans in this loop. They caught it. But they caught it late, with aircraft assigned, which is a thinner margin than the declaration’s language implies.
And the trend since 2023 runs the other direction. Fully autonomous attack drones have been used in the Russian conflict in Ukraine and tested by NATO-backed military contractors.
What happened to the vendor that objected
In March the Department of Defense blacklisted Anthropic over the company’s opposition to its models being used in autonomous weapons systems. Last month a federal judge called that move “unlawful retaliation in violation of the First Amendment.”
So the department punished a supplier for drawing a line on autonomy, lost in court over it, and kept accelerating. Meanwhile extinction-level warnings from AI researchers have pushed AI safety into a national conversation, with calls for regulation and coordinated research “pacing” from leading frontier labs.
If you want a single number to carry out of this, it’s four. Four sources familiar with the episode, none of them the department volunteering it. The near-miss became public because people talked, not because a review process caught a fabrication and published the lesson.