A Hallucinated Cargo Manifest Turns the Pentagon’s AI Push Into a Human-in-the-Loop Problem
A chatbot's invented nuclear-cargo claim about a Chinese vessel nearly triggered a US military operation this spring, and the analyst who trusted it used AI a second time to make the false report...
Military aircraft were already in the air this spring when U.S. officials discovered that the intelligence behind their target was not real. An armed operation against a Chinese-flagged vessel, built on a chatbot’s claim that the ship was carrying components for a nuclear weapons program, was aborted at the last minute. According to multiple sources who spoke to CNN, the report that nearly sent troops to board the ship had been invented from start to finish.
Table Of Content
The near miss, first reported by CNN on September 18 and independently covered by TechCrunch, Gizmodo, and Futurism, happened during the war between the United States and Iran, which the U.S. and Israel have been bombing since February 28, according to Gizmodo’s reporting. The false intelligence claimed a China-flagged ship in the Middle East was carrying nuclear weapons components bound for Iran. It was fabricated from beginning to end, and it got closer to a shooting confrontation with China than anyone in the chain of command realized until the final minutes.
What Actually Happened Aboard the Ship
According to TechCrunch’s reporting, the chain of events started when an analyst at U.S. Special Operations Command queried an AI chatbot to synthesize open-source data with classified signals intelligence about the vessel. The chatbot misidentified the ship’s cargo manifest, inventing a claim that it was hauling components for a nuclear weapons program. The analyst did not stop there. They used the same tool a second time, this round to format those findings into what TechCrunch describes as “an official-looking summary,” which then circulated across command channels.
That summary set the operation in motion. Aircraft launched, and armed personnel prepared to board the vessel. Only at the last minute, with U.S. planes already airborne, did officials realize the intelligence underpinning the operation did not exist. Futurism reports that one source told CNN the intelligence report was “entirely false” and that it “almost started a war” with China.
The Chatbot Still Hasn’t Been Named
One detail is conspicuously missing from every account of the incident: which AI tool actually produced the hallucination. Futurism’s reporting is explicit that the chatbot “has not been identified,” and that it remains unclear whether it was a commercially available product or a system built in-house for government use. That gap matters. A story about one vendor’s model failing under pressure is a story about that vendor’s guardrails. A story about an unnamed tool used inside a classified workflow is a story about whether anyone is tracking which AI systems touch decisions with, in the words of a national security researcher quoted below, “life and death consequences.”
“Not an Isolated Incident”
Gizmodo’s account of the CNN reporting is blunt about how people inside the military are characterizing the episode: a source described the hallucination as part of a broader pattern, not a one-off mistake. Futurism relays a similar warning from an unnamed military source who told CNN, “AI in targeting is definitely something that is ramping up, and there is no real guidance for how having a human in the loop will prevent civilian casualties or fratricide.”
That combination, rapid adoption paired with an admitted absence of guidance, is the actual story here. The Pentagon has been explicit about why it wants AI moving through its decision chain faster. The department has described the technology as delivering what TechCrunch calls “a significant advantage in speeding up its kill chain so commanders can respond in the right time.” Defense Secretary Pete Hegseth has pushed for wider use of AI across the armed forces, and the near miss with the Chinese vessel is the clearest evidence yet of what that push can cost when nothing catches an error before it reaches a commander.
Why the Kill Chain Got Faster
The military’s “speed of thought” framing predates this incident by years, but the war with Iran gave it a concrete test. Craig Jones, a lecturer at Newcastle University and author of The War Lawyers: The United States, Israel, and Juridical Warfare, told Fortune that the U.S. Air Force has used “speed of thought” as a decision-making benchmark for years, and that intelligence-to-strike cycles which once took up to six months during World War II and Vietnam have been compressed dramatically by tools that fuse together, in his words, “terabytes and terabytes and terabytes of data, everything from aerial imagery, human intelligence, internet intelligence, mobile phone tracking, anything and everything.”
Jones argues the compression is not hypothetical. He told Fortune the U.S.-Israeli strikes on Iran, which resulted in the death of Ayatollah Ali Khamenei, “would have been impossible, or almost impossible, to do in that way,” adding that “the speed it was carried out, and the magnitude and the volume of the strikes, I think, are AI-enabled.” Amir Husain, coauthor of Hyperwar: Conflict and Competition in the AI Century, told Fortune that AI now plays a role across the military’s entire observe-orient-decide-act loop, including autonomous drones that must act without human guidance when signals are jammed. Husain was just as clear about where he thinks responsibility still belongs: “The laws of armed conflict require us to blame the person. The person has to be accountable no matter what level of automation is used in the battlefield.”
The same reporting notes that Claude, Anthropic’s AI model, was reportedly used by the U.S. military in its attack on Iran, according to the Wall Street Journal, despite President Trump telling federal agencies and military contractors to cease doing business with Anthropic. That order followed a public falling-out over how Claude was being used inside the Pentagon, and it was later found to be illegal: a federal judge ruled the blacklisting was retaliation for Anthropic’s criticism of the Department of Defense. Nothing in the current reporting ties Claude, or any other named model, to the Chinese vessel hallucination specifically. But the episode lands in the middle of a defense AI market where vendor relationships have already proven to be a live legal and political fight, not a settled one.
The Accountability Gap
Jake Steckler, a research scholar at the AI governance nonprofit GovAI and a veteran U.S. Army aviation officer, gave TechCrunch the most direct assessment of what the incident should change. “It’s important for service members to understand the uncertainty inherent to LLMs,” Steckler said. “But it’s especially critical for any decisions that could lead to use of force, like targeting, intelligence analysis, or operational planning. There are life and death consequences for those decisions.”
Steckler was careful not to frame the near miss as a reason to pull AI out of military workflows entirely. “These tools can be useful in the right contexts and with the right safeguards in place,” he told TechCrunch. His warning was about what happens if speed keeps winning over verification: “Prioritizing adoption speed over all else will likely lead to incidents that only make service members lose trust in these systems, which ultimately is only going to slow adoption.”
A Second AI Query Isn’t Verification
The detail that should worry anyone building or buying AI tools for high-stakes decisions is not that a chatbot hallucinated. Every model does that under the wrong conditions. It is that the analyst’s response to a plausible-looking answer was to ask the same tool to dress it up for circulation, rather than to check it against an independent source. A hallucination caught before it leaves someone’s screen is a bug. A hallucination reformatted into “an official-looking summary” and routed up a chain of command is a process failure, and it is the kind of failure that shows up regardless of which vendor’s model sits underneath it.
A Recurring Pattern in How the Pentagon Buys and Deploys AI
This is not the first time this year that the mechanics of how the Pentagon adopts AI, rather than the AI itself, has been the story. In August, a Pentagon memo directed up to $244 million to Palantir without competitive bidding, a decision that raised its own conflict-of-interest questions. Between a no-bid contract awarded outside normal oversight and a chatbot-generated intelligence report that reached commanders without independent verification, the common thread is the same: procurement and deployment decisions that move faster than the checks meant to catch them.
The fix Steckler is describing already exists as a body of practical technique, not just a policy aspiration. Retrieval-scope checks, confidence scoring on individual claims, and a hard requirement that a system flag when it does not have enough grounding to answer are the kind of layered defenses this site has walked through building from scratch. None of that is exotic. It is the difference between a chatbot that answers every question with equal confidence and one that says, correctly, that it does not know.
What Would Actually Have to Change
Nobody involved has said whether the Pentagon plans to identify the chatbot in question, investigate how a fabricated report cleared multiple layers of review, or add a mandatory independent-verification step before an AI-assisted intelligence summary can reach a commander. Until one of those things happens, the honest read of this incident is simple: a hallucinated cargo manifest got as close to an armed confrontation with China as an aircraft already in the air, and the process that let it get there is still in place today.








No Comment! Be the first one.