Bitcoin Red Team, an AI-assisted security campaign aimed at Bitcoin projects, says it generated 6,700 findings across 425 projects in its first 55 hours. It labeled 1,029 of those high or critical.
What it hasn’t said is how many were real.
That gap is the whole story. The Aug. 6 update is a measurement of how much material entered a triage pipeline, not of how much software got safer. The campaign’s effect on security remains unreported.
The missing numbers are the ones that matter
The thread published no audit-ready definitions for its severity labels and no denominators behind them. No case-level outcomes. No aggregate false-positive rate. No fix rate.
Without those fields you can’t calculate how many alerts became confirmed vulnerabilities, how many maintainers rejected or downgraded, or how many produced a patch. You’re left with a headline count and a severity split, both of them self-assessed.
Which doesn’t make the 55 hours meaningless. It shows something worth taking seriously: an AI system can fill a review pipeline at the scale of an entire set of projects, fast. Expert prompting, reproduction, disclosure and maintainer response were still required at every stage after that.
Two snapshots, and what changed between them
The campaign published progress twice. At 27.5 hours it covered 390 projects and 4,962 findings. By 55 hours the project count had risen by 35 and the finding count by 1,738.
The later thread put high-or-critical findings at 15.4% of the total, and clarified that three of the 24 reported participants were bots.
Worth noticing how the accounting shifted: the earlier post separated critical from high, the later one combined them. Both sets of figures reflect campaign assessments. Maintainer-confirmed exploitability and remediation outcomes need separate evidence, and that evidence isn’t in either post.
The models searched. People decided.
Rob Hamilton described Kimi K3 as handling the heavy analysis, with GPT Sol, Fable/Opus and GLM 5.2 supporting the documentation. He said OpenAI’s Cyber Harness covered selected components he considered load-bearing.
A day later, Hamilton wrote that subject-matter experts could change an assessment with one or two sentences of context or a small block of code. In the examples he described, that input pushed middling concerns into high or critical territory.
He also named the bottlenecks: operations, disclosure handoff and triage. Not model capacity.
In his account, models searched broadly while specialists shaped prompts, interpreted output, attempted reproduction and decided which reports were ready for disclosure. That division of labor is what makes this a human-AI review system rather than a scanner with a press release.
$20,000 and 150 repositories
On Aug. 3, Hamilton said the effort had spent over $10,000 scanning over 100 repositories, and had immediately disclosed critical findings when a proof of concept demonstrated exploitability. On Aug. 4, he reported about $20,000 in spending, more than a dozen disclosures and 150 repositories scanned.
Scanning kept expanding. Outreach, handoff and triage stayed described as active operational constraints. The snapshots offer no comparable disclosure denominator at 55 hours, so you can’t compare how fast findings arrived against how fast they got resolved.
Hamilton later identified the separate Coldcard incident as a catalyst for the wider campaign. The campaign record attributes no discovery of the Coldcard flaw to this sprint.
Most projects have nobody to email
Here’s the finding I’d argue is the most useful thing the sprint produced. In the 55-hour update, Bitcoin Red Team reported that 19.5% of scanned projects had a SECURITY.md file and 13.1% had an email in it.
The thread omitted the project corpus, how the denominator should be read, and the measurement method, so those percentages describe the campaign’s scan and nothing broader. Still: if you flood a pipeline with findings and the projects on the other end have no documented way to receive them, volume isn’t your constraint.
The critics have the same evidence problem
The developer known as Calle said most critical reports were quickly verified by project owners. The post supplied no denominator, no verified-report count, no rejection count and no patch status, which leaves the breadth and outcome of that verification unresolved.
On the other side, JW Weatherman argued publicly that the campaign couldn’t triage its own output. His post identified no campaign-linked issue, patch or advisory, so it’s criticism without a measurable failure rate.
Both claims run into the same wall. The disposition data isn’t public, so neither can be checked.
What a real accounting would look like
Separate the findings that were reproduced, acknowledged, downgraded, rejected and fixed. Publish a definition and a denominator for each rate. That breakdown would show how much of the campaign’s volume became actionable work, and it’s the difference between a security result and a throughput result.
For now, 6,700 is a count of campaign-labeled findings and triage candidates. The sprint proved machine-assisted review is fast. Its lasting value depends entirely on the share that experts can validate, disclose and turn into patches, and that’s a number the campaign hasn’t published yet.