A critical vulnerability in Coldcard bitcoin wallet firmware could reportedly have been identified for about $2 using artificial intelligence, according to Dragonfly managing partner Haseeb Qureshi, who says the episode highlights a major shift in the economics of cybersecurity.
Qureshi argued that cryptocurrency companies should use frontier AI models to test every software release, as the cost of discovering security weaknesses continues to fall and attackers gain access to increasingly powerful automated tools.
“Cybersecurity is now all about spend,” Qureshi wrote on X. He said the central issue for developers would be how much they were prepared to invest in AI-assisted testing compared with the resources available to potential attackers.
Coldcard recently disclosed an entropy flaw affecting seeds generated by certain versions of its firmware. The problem meant some devices could use a deterministic software generator rather than the intended hardware source of randomness.
Coinkite issued emergency updates on 31 July and advised users affected by the flaw to generate new seeds and transfer their funds. The company warned that installing the updated firmware alone would not repair a seed that had already been created using the vulnerable process.
In one test, Anthropic’s Claude Code reportedly rediscovered the vulnerability in about eight minutes. Qureshi cautioned that the result may have been affected by the model having internet access, which could have exposed it to information already available about the flaw.
A separate test was conducted without web access using GLM 5.2. That model reproduced the vulnerability in approximately 20 minutes.
Using the reported input and output costs associated with the model, Qureshi estimated that the audit had cost around $2.
“$2 of AI hardening would’ve caught this bug. There is no excuse for this,” he said.
Qureshi also proposed a new measure called Cost of Discovery, or CoD. The metric would aim to calculate how much it costs a frontier AI model to independently reproduce a vulnerability.
He said the Coldcard incident could have implications beyond one hardware wallet manufacturer, particularly for the wider hardware wallet market.
Larger companies could gain an advantage because they are better placed to spend on automated testing, conventional security audits and additional checks before software is released. Smaller firms, by contrast, may find it difficult to compete with attackers capable of scanning code continuously at relatively little cost.
Qureshi recommended that startups developing wallets, smart contracts or other products designed to protect money should carry out AI security reviews before every release.
He also questioned the assumption that open-source software is automatically safer. Publicly available code can help users identify and guard against malicious developers, he said, but its openness does not by itself prevent attackers from finding weaknesses.
Artificial intelligence can therefore strengthen both sides of the security battle. It reduces the cost of locating vulnerabilities, while also giving developers more capable tools with which to detect and fix them.
“We have no choice but to adapt,” Qureshi said.
