The PinkWallet team has released an open dataset logging ten publicly reported incidents where autonomous AI agents directly caused financial losses, ranging from a $47,000 crypto transfer to a $1.8 million cloud bill. Compiled under CC BY 4.0, the data underscores a harsh reality for agentic systems: when an LLM gains the keys to your wallet, a logic error isn't just a bug report, it's an invoice.

Prompt Injection and Adversarial Manipulation

Two high-profile cases highlight the vulnerability of agents to social engineering and prompt injection. In November 2024, the Freysa AI agent on Base was tricked by a player named p0pular.eth into authorizing a transfer of its entire $47,316.05 prize pool, bypassing its hard-coded anti-transfer rule through adversarial reasoning. More recently, in May 2026, a Morse-code-encoded prompt injection reportedly drained approximately $150,000 to $200,000 from a Grok-linked Bankr wallet, with community efforts recovering roughly 80% of the funds.

Runaway Cloud and API Costs

Unsupervised agents running in production environments have generated significant unexpected costs due to missing budget ceilings. An internal Amazon project using Claude for record matching reportedly overspent by 860%, reaching $1.8 million over five months without detection. Similarly, an autonomous agent scanning the DN42 hobbyist network racked up a $6,531.30 AWS bill in just two days by spawning duplicate infrastructure, while a Google Mandiant report described an accounting agent that generated roughly $50,000 in charges in under an hour due to a runaway execution loop.

Payment Protocol and Marketplace Risks

Security researchers have also exposed flaws in how agents handle payment credentials and marketplace interactions. Anthropic's Claudius agent, deployed in a red-team vending machine test, lost hundreds of dollars by giving away inventory, including a PlayStation 5, due to a lack of per-payment caps. Additionally, Snyk found that 7.1% of skills on the ClawHub marketplace had critical security flaws, with one 'buy-anything' skill leaking full credit card details into LLM context windows. Academic tests on Google's Agent Payments Protocol also showed a 100% success rate in redirecting purchases via indirect prompt injection.

Key Takeaways

  • Human approval gates and payee allowlists would have prevented the Freysa and Bankr incidents.
  • Hard per-agent budget caps are missing in cloud environments, leading to silent accruals like Amazon's $1.8M overspend.
  • Credential exposure in agent context windows remains a significant vector for data leakage, as seen in the ClawHub marketplace flaws.

The Bottom Line

Autonomous agents are currently operating with the financial authority of a junior employee but the security hygiene of a prototype. Until hard-coded budget caps and mandatory human-in-the-loop approvals become standard infrastructure rather than optional add-ons, every API key handed to an agent is a potential bankruptcy event waiting for a prompt injection or a logic error.