Giving an AI agent a wallet, and keeping it inside a limit
An AI agent wallet is an ordinary crypto account whose private key happens to sit inside a program rather than in a person's pocket, and the moment you fund it, the agent can spend it. The only spending limit that survives a bad day is one enforced somewhere the agent cannot reach: the balance of the account it signs from, or a wallet level spend permission that a contract checks before it moves anything. A cap written into a system prompt is a suggestion. A cap in your client library is a real defence against your own bugs and a thin one against an attacker who can write into the model's context, because that attacker is executing inside the same process as the cap. Fund a dedicated account with a sum you would shrug at, put a per payment cap under that balance, and assume every payment is final, because a signed stablecoin authorisation is.
What does it actually mean to let an AI agent pay?
In an x402 flow the agent makes a normal HTTP request, gets a 402 Payment Required back with machine readable terms, signs a stablecoin transfer authorisation, and retries with that signature attached. No card, no account, no API key. On EVM chains the signature is an ERC-3009 transferWithAuthorization message, a standard whose stated purpose is to let a user "delegate the gas payment to someone else" while signing for the transfer itself.
So the agent never sends a transaction. It signs a message that authorises one. Somebody else broadcasts it. That somebody is the facilitator, and on this origin the arrangement is stated plainly: "It holds no funds: a settlement moves USDC from the payer to the merchant's advertised address in one transaction that the payer signed; the facilitator broadcasts it and pays the gas." The merchant does not hold your key either. The x402 FAQ notes that sellers "never hold the buyer's key; they only verify signatures."
That is good for key hygiene and bad for anyone hoping a middleman will claw a payment back. No middleman holds the money.
Where can a spending limit be enforced?
There are three places, and they are not equally strong.
The wallet layer is the real one. An account holding 20 USDC cannot spend 21. Better than a bare balance is a smart account with a signed permission: Coinbase's Spend Permissions documentation describes it as designating a trusted spender, and says that "After you sign the permission, the spender can initiate token spending within the limits you define. You can define limits based on token, time period, and amount." The project's README is explicit that this design "does not enable apps to make arbitrary external calls from user accounts". ERC-7715 generalises the same idea into a wallet RPC method, defining wallet_requestExecutionPermissions "for DApp to request a Wallet to grant permissions in order to execute transactions on the user's behalf".
The client layer is your agent's own code. The reference x402 client ships with this on by default: "By default, the x402 client only pays recognized USD-pegged assets (e.g. USDC) and caps each payment at $1." The docs add a detail worth reading twice: "Spend controls run before any custom policies and before the payment payload is signed." A payment that fails the cap is never signed, so there is nothing to broadcast.
The merchant layer is the price in the 402 response. It is not your limit at all. A stranger chose it, and a hostile merchant can change it between one request and the next.
Which of these can an attacker who controls the prompt bypass?
Both of the top two, and neither is a close call.
OWASP catalogues prompt injection as LLM01 in its Top 10 for LLM Applications, and is blunt about the prognosis: "Prompt injection vulnerabilities are possible due to the nature of generative AI. Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection." The dangerous variety for a paying agent is the indirect kind. OWASP: "Indirect prompt injections occur when an LLM accepts input from external sources, such as websites or files."
An agent with a wallet reads the web, which is usually the whole point of it, so the untrusted text it summarises can carry instructions and the model has no reliable way to tell an instruction from a quotation. If your cap lives in a config value the agent can rewrite, or in a prompt it can be talked out of, the attack is a paragraph of text. If your cap lives in the account balance, the chain simply refuses.
Two OWASP mitigations apply directly. "Restrict the model's access privileges to the minimum necessary for its intended operations" means a wallet holding this week's budget and nothing else. "Require human approval for high-risk actions" is why the x402 client documentation points at lifecycle hooks rather than caps when you want a person in the loop.
A worked example: one funded key, two caps
The SDK cap is per payment. Two hundred calls at five cents each is ten dollars, and every one passes a five cent cap. Pair it with a running total.
import json, pathlib, time
from decimal import Decimal
from x402 import x402Client, x402ClientConfig, SchemeRegistration
from x402.mechanisms.evm.exact import ExactEvmScheme
PER_CALL_CAP = "$0.05" # no single payment may exceed this
DAILY_CAP = Decimal("2.00") # total dollars this agent may spend in a day
BOOK = pathlib.Path("spend.json")
# signer holds the agent's key. It is NOT your treasury key.
client = x402Client.from_config(
x402ClientConfig(
schemes=[SchemeRegistration(network="eip155:8453", client=ExactEvmScheme(signer))],
spend_controls={"max_amount_per_payment": PER_CALL_CAP},
)
)
def _ledger() -> dict:
return json.loads(BOOK.read_text()) if BOOK.exists() else {}
def budget_left() -> Decimal:
return DAILY_CAP - Decimal(_ledger().get(time.strftime("%Y-%m-%d"), "0"))
def record(amount_usd: str) -> None:
book, day = _ledger(), time.strftime("%Y-%m-%d")
book[day] = str(Decimal(book.get(day, "0")) + Decimal(amount_usd))
BOOK.write_text(json.dumps(book))Test the whole path against a live merchant before you point it at a real one:
curl -i https://pay.neuronto.com/echo/echo answers 402, charges $0.001 on Base and returns it in the same request, with the gas covered by the origin. A replayed payment is answered from the record instead of charged twice, so a retry loop will not quietly bill you again while you debug.
Now read the code again with the attack in mind. DAILY_CAP, PER_CALL_CAP and spend.json all sit inside the process the agent controls. They stop runaway loops and off by one errors, which is most of what goes wrong. They do not stop an agent persuaded to call record with the wrong number. Only the balance behind signer does that.
What x402 does and does not protect against
x402 is a payment rail. It moves an agreed amount from a payer to a merchant and proves it happened. It is not an authorisation policy engine, and nothing in the protocol asks whether this agent should be buying this thing right now.
What it gives you: a signature scoped to one amount, one recipient and one time window, rather than a reusable credential. ERC-3009 uses a random 32 byte nonce per authorisation, so a payload cannot be replayed once spent. A facilitator that simulates the transfer immediately before broadcasting will fail cleanly rather than burn gas on a payer whose balance moved. Settlement is idempotent per body on this origin, so a retry returns the recorded outcome instead of paying twice.
What it does not give you: a refund. The x402 FAQ calls the exact scheme "a push payment" that is "irreversible once executed", and the only remedies it lists are the merchant voluntarily sending money back, or a different scheme built on escrow. There is no dispute process, no chargeback and no issuer to call. It also does not give you intent: the protocol has no notion of what the purchase was for. Google's AP2 work tries to fill that gap with signed mandates that specify "price limits, timing, and other conditions", but AP2 sits above the rail, not inside it.
And the rail carries its own sharp edges. ERC-3009's security notes warn that "It is possible for an attacker watching the transaction pool to extract the transfer authorization and front-run the transferWithAuthorization call to execute the transfer without invoking the wrapper function", which is why contracts that receive such payments should use receiveWithAuthorization instead.
Before you fund it
- Use a fresh account. Never the key that holds your treasury, and never a key reused for anything else.
- Fund it with a number you would accept losing outright today, not a number you expect to spend this month.
- Top it up on a schedule rather than granting an allowance you forget about. The refill interval is the real limit.
- If the wallet supports a signed spend permission with an amount and a period, use it, and record how to revoke it before you need to.
- Set the client cap below the largest legitimate single purchase, not above it.
- Keep a separate running total per day or per run, and make exceeding it a hard stop rather than a warning.
- Log every 402 the agent accepted, including the recipient address. You cannot reverse a payment, so the log is the whole of your forensics.
- Decide in advance which purchases require a person, and put that check in code rather than in the prompt.
- Test against a refunding merchant before anything that matters.
Common questions
Does an AI agent need a crypto wallet to pay for things?
For x402 it needs a signing key and a funded address. The x402 documentation treats the wallet as "both a payment mechanism and a form of unique identity for buyers and sellers", so the address doubles as the agent's identifier to merchants.
Can I set a spending limit that the agent cannot change?
Yes, but not in the agent. Put it in the balance of the account or in a signed permission that a contract enforces. Anything inside the agent's own process is changeable by anything that gets into that process.
Is a per payment cap enough?
No. It bounds one purchase. Volume is a separate risk needing its own counter, with a funding ceiling underneath both.
Can the facilitator take my money or freeze it?
It never holds it. A settlement is one transaction the payer signed, moving USDC from payer to the merchant's advertised address, broadcast by the facilitator. It can decline to broadcast, which stops a payment. It cannot redirect one.
What happens if the agent pays and gets nothing back?
You have paid. Serve after settle is the merchant's discipline, not a guarantee to you. Keep a per merchant spend total so one bad counterparty cannot drain a day's budget.
Sources
- x402 documentation, FAQ. https://docs.x402.org/faq
- x402 documentation, Quickstart for Buyers (Spend Controls). https://docs.x402.org/getting-started/quickstart-for-buyers
- x402 documentation, Exact scheme. https://docs.x402.org/schemes/exact
- x402 documentation, Wallet. https://docs.x402.org/core-concepts/wallet
- ERC-3009: Transfer With Authorization, Ethereum Improvement Proposals. https://eips.ethereum.org/EIPS/eip-3009
- ERC-7715: Request Permissions from Wallets, Ethereum Improvement Proposals. https://eips.ethereum.org/EIPS/eip-7715
- OWASP Top 10 for LLM Applications, LLM01:2025 Prompt Injection. https://owasp.org/www-project-top-10-for-large-language-model-applications/
- Coinbase Developer Documentation, Spend Permissions. https://docs.cdp.coinbase.com/wallets/using-wallets/spend-permissions
- coinbase/spend-permissions, README. https://github.com/coinbase/spend-permissions
- Google Cloud, Announcing the Agent Payments Protocol (AP2). https://cloud.google.com/blog/products/ai-machine-learning/announcing-agents-to-payments-ap2-protocol
- Neuronto Payments, developer reference. https://pay.neuronto.com/developers
- Neuronto Payments, integrate. https://pay.neuronto.com/integrate