AI-assisted safety marketing campaign targeted on the Bitcoin ecosystem, Bitcoin Pink Workforce, stated it generated 6,700 findings throughout 425 initiatives in its first 55 hours. The marketing campaign labeled 1,029 of them excessive or important.
The Aug. 6 replace measures how a lot materials entered a safety triage pipeline, and its impact on software program safety stays unreported.
The retrieved thread omitted audit-ready definitions and denominators for the severity counts, in addition to case-level outcomes, an mixture false-positive fee, and a repair fee.
These lacking fields stop a calculation of what number of alerts turned confirmed vulnerabilities, what number of maintainers rejected or downgraded, and what number of led to patches.
The primary 55 hours nonetheless reveal a consequential functionality, noting how AI techniques can fill an ecosystem-scale evaluate pipeline rapidly. Skilled prompting, copy, disclosure, and maintainer response remained vital at each later stage.
What the marketing campaign numbers measure
The marketing campaign revealed two snapshots as its roster and workload expanded:
| Elapsed time | Tasks | Whole findings | Reported severity | Members |
|---|---|---|---|---|
| 27.5 hours | 390 | 4,962 | 85 important; 635 excessive | 16 |
| 55 hours | 425 | 6,700 | 1,029 excessive or important | 24 reported, together with three bots |
The 27.5-hour replace coated 390 initiatives and 4,962 findings. By the 55-hour mark, the challenge rely had risen by 35 and the discovering rely by 1,738. The later thread put high-or-critical findings at 15.4% of the entire and clarified that three of the 24 reported members have been bots.
The sooner publish separated important and excessive findings, whereas the later one mixed them, with each units of figures reflecting marketing campaign assessments. Maintainer-confirmed exploitability and remediation outcomes require separate proof.
Rob Hamilton described Kimi K3 as dealing with the heavy evaluation, with GPT Sol, Fable/Opus, and GLM 5.2 supporting the documentation. He stated OpenAI’s Cyber Harness coated chosen elements he thought-about load-bearing.
A day later, Hamilton wrote that subject-matter consultants may change an evaluation with one or two sentences of context or a small block of code. In examples he described, that enter pushed middling issues into excessive or important territory. He additionally recognized operations, disclosure handoff, and triage as bottlenecks.
In Hamilton’s account, fashions searched broadly whereas specialists formed prompts, interpreted output, tried copy, and determined which experiences have been prepared for disclosure. That division of labor makes the marketing campaign a human-AI evaluate system.
The developer referred to as Calle stated most crucial experiences have been rapidly verified by challenge homeowners. The publish equipped no denominator, verified-report rely, rejection rely, or patch standing, leaving the breadth and final result of that verification unresolved.
Outreach and outcomes outline the safety worth
Within the 55-hour replace, Bitcoin Pink Workforce reported that 19.5% of scanned initiatives had a SECURITY.md file and 13.1% had an e-mail there. The retrieved thread omitted the challenge corpus, denominator interpretation, and measurement technique, so the odds solely describe the marketing campaign’s scan.
On Aug. 3, Hamilton stated the hassle had spent over $10,000 scanning over 100 repositories and had instantly disclosed important findings when a proof of idea demonstrated exploitability. On Aug. 4, he reported about $20,000 in spending, greater than a dozen disclosures and 150 repositories scanned.
Scanning continued to develop, whereas the marketing campaign described outreach, handoff and triage as energetic operational constraints. The revealed snapshots supply no comparable disclosure denominator at 55 hours, so they can not set up the relative pace of scanning and backbone.
Hamilton later recognized the separate Coldcard incident as a catalyst for the broader marketing campaign. The marketing campaign report attributes no discovery of the Coldcard flaw to this dash.
A helpful public accounting would separate findings that have been reproduced, acknowledged, downgraded, rejected, and glued, with definitions and denominators for every fee. That breakdown would present how a lot of the marketing campaign’s quantity turned actionable safety work.
A public critic, JW Weatherman, argued that the marketing campaign couldn’t triage its output. His publish recognized no campaign-linked challenge, patch, or advisory, so it provides criticism with out a measurable failure fee. The marketing campaign’s lacking disposition knowledge leaves the underlying query open.
For now, 6,700 represents campaign-labeled findings and triage candidates. The dash demonstrated the pace of machine-assisted evaluate. Its lasting safety worth depends upon the share that consultants can validate, disclose, and convert into fixes.

