Neuronto PaymentsAccept payments

How to charge per API request, and what an API call actually costs

An API call costs two very different things, and mixing them up is where most API pricing goes wrong. Serving an ordinary request costs the provider a small fraction of a cent. Cloudflare Workers bills $0.30 per additional million requests past the 10 million included each month on its Standard plan. AWS Lambda charges $0.20 per million requests plus compute. Amazon API Gateway charges $1.00 per million calls for HTTP APIs and $3.50 per million for REST APIs. That is roughly $0.0000003 to $0.0000035 per call. A call that runs a model behind it costs about a thousand times more. At Anthropic's published rate for Claude Haiku 4.5, $1 per million input tokens and $5 per million output tokens, a single request with 1,000 input tokens and 500 output tokens costs $0.0035. So "what does an API call cost" has no single answer: the serving cost and the work cost differ by four orders of magnitude, and the price you charge is set by what the answer is worth to the buyer, not by either. What does put a hard floor under your price is payment overhead. A card charge carries a fixed fee per successful transaction, so a $0.01 sale is not worth collecting on cards. A stablecoin settlement over HTTP does not have that floor, which is why per-request pricing has only recently become practical at cent and sub-cent amounts.

The three numbers people call "the cost of an API call"

What is being measuredVerified figureSource
Cost to route the request$1.00 per million (HTTP API), $3.50 per million (REST API)Amazon API Gateway pricing
Cost to execute a function$0.20 per million requests, plus GB-secondsAWS Lambda pricing
Cost to serve from the edge$0.30 per additional million past 10 million includedCloudflare Workers pricing
Cost of the work, if it is inference$0.0035 for a 1,000-in / 500-out call on Claude Haiku 4.5Anthropic pricing
Cost to collect the money on cardsa percentage plus a fixed fee per successful charge (1.5% + EUR 0.25 for standard EEA cards on Stripe's published European rates)Stripe pricing
Cost to collect the money over HTTP$0 per verification; $0.00211 estimated per on-chain settlement on BaseNeuronto Payments pricing

If your API wraps a model, cost per call is dominated by tokens and per-call pricing tracks it exactly. If it serves cached data, your marginal cost is near zero and per-call pricing is a positioning choice, not cost recovery.

Four ways to get paid, and what each one actually costs you

Four API pricing models plotted by buyer friction against revenue predictability Buyer friction against revenue predictability Per request Freemium Prepaid credits Subscription low friction high friction low high Horizontal: what a first-time buyer must do before the first call succeeds. Vertical: how well this month's revenue predicts next month's.
ModelBuyer has toYou carryBreaks when
Per requestNothing but pay for the callNo dunning, no invoices, no seat management; revenue tracks usage exactlyPer-payment overhead approaches the price; finance wants a forecast; the buyer's procurement needs an invoice
SubscriptionSign up, pick a tier, give a cardChurn, tier design, support for buyers who outgrow or underuse a tierUsage is spiky; the buyer only needs you once; a heavy user destroys your margin inside a flat tier
Prepaid creditsCommit money before knowing the valueDeferred revenue accounting, expiry policy, refund requests, balance support ticketsBuyers run out mid-job, or overbuy and resent it; unused balances become a liability and a support queue
FreemiumNothing up frontThe whole free tier's cost, plus abuse controls, plus a conversion funnel that may never convertFree usage is expensive (inference is), or the free tier is good enough that nobody upgrades

None of these is correct in general. A weather lookup at $0.0001 per call cannot be sold per request, because the payment costs more than the data. A one-off document conversion cannot be sold as a subscription, because nobody subscribes for one job.

How charging per request actually works over HTTP

HTTP has had a status code reserved for this since the beginning. RFC 9110, section 15.5.3, says in full: "The 402 (Payment Required) status code is reserved for future use." It sat unused because a machine had no way to pay another machine without a human first filling in a form.

x402 uses that code. The x402 documentation describes it as "the open payment standard that enables services to charge for access to their APIs and content directly over HTTP", "built around the HTTP 402 Payment Required status code", which "allows clients to programmatically pay for resources without accounts, sessions, or credential management". The flow has four steps:

  • An unpaid request arrives. Your server answers 402 with machine-readable terms: price, asset, network, and the address to pay.
  • The client signs a stablecoin transfer authorisation for exactly that amount. It does not broadcast anything.
  • The client retries the same request, carrying the signed payload in a header.
  • Your server hands the payload to a facilitator, which verifies the signature and settles the transfer on-chain, then returns the resource.

The division of labour is the part worth understanding. Your server never holds a private key, never talks to a blockchain, and never stores the buyer's identity. The facilitator does the chain work and pays the gas. The x402 documentation is explicit that the standard itself adds nothing on top: "x402 as a standard has 0 fees built in." Whatever it costs is the facilitator's price plus network gas.

A worked example

Point an x402 server SDK at a facilitator and advertise a wallet you control. That is the whole integration. Node and Express:

js
import express from "express";
import { paymentMiddleware, x402ResourceServer } from "@x402/express";
import { ExactEvmScheme } from "@x402/evm/exact/server";
import { HTTPFacilitatorClient } from "@x402/core/server";

const server = new x402ResourceServer(
  new HTTPFacilitatorClient({ url: "https://pay.neuronto.com" }),
).register("eip155:8453", new ExactEvmScheme());

const app = express();
app.use(paymentMiddleware({
  "GET /premium": {
    accepts: { scheme: "exact", price: "$0.01", network: "eip155:8453", payTo: "0xYourWallet" },
    description: "One premium answer",
    mimeType: "application/json",
  },
}, server));
app.get("/premium", (req, res) => res.json({ ok: true }));
app.listen(3000);

Python and FastAPI:

python
from fastapi import FastAPI, Request
from x402 import x402ResourceServer
from x402.http import FacilitatorConfig, HTTPFacilitatorClient
from x402.http.middleware.fastapi import payment_middleware
from x402.mechanisms.evm.exact import register_exact_evm_server

server = x402ResourceServer(HTTPFacilitatorClient(FacilitatorConfig(url="https://pay.neuronto.com")))
register_exact_evm_server(server, "eip155:8453")

middleware = payment_middleware({
    "GET /premium": {
        "accepts": {"scheme": "exact", "payTo": "0xYourWallet", "price": "$0.01",
                    "network": "eip155:8453"},
        "description": "One premium answer",
        "mimeType": "application/json",
    }
}, server)

app = FastAPI()


@app.middleware("http")
async def x402(request: Request, call_next):
    return await middleware(request, call_next)


@app.get("/premium")
def premium():
    return {"ok": True}

The unpaid half of the exchange is visible without writing any code:

bash
curl -i https://pay.neuronto.com/echo

That endpoint charges $0.001 and returns it in the same request, so a client can be tested end to end without losing money.

Setting the number: how to determine the price of an API call

Work upward from three floors and stop at the ceiling.

  • Floor one, marginal cost. Tokens, compute, third-party data you resell. For the Haiku-shaped call above, $0.0035.
  • Floor two, payment overhead. Every settled payment costs something. Neuronto Payments publishes its settlement price as network gas cost multiplied by 1.3, rounded up to a hundredth of a credit, with a credit defined as $0.001; on Base that currently estimates to $0.00211 per settlement, verification is $0, and any change is announced 14 days ahead. That number matters more than it looks: if you charge $0.001 per call and settle every call individually, the settlement costs roughly twice your revenue. Per-request settlement wants prices around a cent and up, or batching.
  • Floor three, the cost of being wrong. Support, abuse, the one customer who sends ten million calls.
  • Ceiling, substitution. What the buyer pays today to get the same answer another way, including the cost of building it themselves.

Price near the ceiling, not near the floor. Cost tells you when to refuse a deal; it does not tell you what to charge.

When per-request billing is the wrong choice

High-frequency, tiny calls. If you sell a $0.0002 lookup at a million calls a day, per-call settlement is arithmetic nonsense: the fee is an order of magnitude above the price. Sell a bucket, a subscription, or batch the settlements. Charging per request is not free just because the protocol is open.

Anything that needs a refund. Per-request stablecoin settlement is final. There is no chargeback, which is the selling point from the merchant's side and the disqualifier from the buyer's. If your product can fail in ways that deserve money back, you need a system that can reverse a payment, and a card processor is the right tool despite the fixed fee.

Enterprise procurement. A company that needs a signed order form, net 30 terms, a VAT invoice, a security review and a named account manager cannot buy per request, whatever the API does. The blocker is the finance department, not the protocol. Per-request billing is for self-serve buyers and agents.

Buyers who cannot hold stablecoins. An x402 route is payable only by a client with a funded wallet on a supported network. If your buyer is a person with a corporate card, this is a worse checkout than a payment link, not a better one. It becomes the better checkout when the buyer is software.

Revenue you need to forecast. Usage-based revenue is honest and volatile. If you are hiring against a plan or raising on ARR, a pure per-call model hands you a number that moves with someone else's traffic.

The honest summary: per-request payment is best when the buyer is an autonomous client, the price is at least a cent, the transaction is final, and onboarding friction is costing you sales. Outside that, another model in the table is better.

FAQ

How much does an API call cost?

To serve, between roughly $0.0000003 and $0.0000035 on common serverless platforms (Cloudflare Workers, AWS Lambda, Amazon API Gateway). To run, $0.0035 for a small Claude Haiku 4.5 call at published rates. To buy, whatever the provider charges, which is a value decision rather than a cost pass-through.

What is an API fee?

Two unrelated things travel under that name. One is what a provider charges you per call or per month to use their API. The other is what a payment provider takes to move the money, which for cards is a percentage plus a fixed amount per successful charge, and for an x402 settlement is the network gas plus the facilitator's margin.

Cost per API call for Claude or AWS specifically?

Anthropic publishes per-million-token rates, so the per-call cost is arithmetic on your own token counts: $1 and $5 per million (input and output) on Claude Haiku 4.5; Claude Sonnet 5 is $2 and $10 per million. AWS publishes per-million-request rates for API Gateway and Lambda. For any other provider, read their own pricing page rather than a third-party comparison.

How do I monetize an API without building billing?

Put a payment middleware in front of the routes you want to charge for, declare a price per route, and point it at a facilitator. There is no account system, no metering database, no invoicing, and no key rotation, because the payment arrives with the request.

Can you monetize MCP servers?

Yes, by the same mechanism, because an MCP server is served over HTTP. The tool call that costs you money is the one you put behind a price.

Is pay per API call better than a subscription?

It is better at acquisition and worse at forecasting. Per-call pricing removes every step between a stranger and a working request. Subscriptions give you revenue you can plan around. Many APIs end up with both: a per-call route for first contact, a committed plan for anyone who stays.

Sources