
Pay-per-inference for AI agents: How BlockRun and Incarna use Amazon Bedrock AgentCore payments
Quick Answer
Amazon Bedrock AgentCore payments enables AI agents to autonomously pay for model inferences, streamlining the process for Incarna and BlockRun.
Quick Take
This managed service allows agents to handle microtransactions efficiently, ensuring spending limits are enforced at the infrastructure level, thus reducing integration time from months to days.
Key Points
- Incarna's agents use AgentCore payments for real-time model inference payments.
- BlockRun serves over 90 models, charging per inference request.
- AgentCore payments supports x402 payment protocols for efficient transactions.
- Spending limits are enforced outside the model to prevent overspending.
- Payments are settled in USDC, ensuring transaction verifiability on-chain.
DeepSignal Analysis
What happened
Amazon Bedrock AgentCore payments allows AI agents to autonomously handle payments for model inferences, streamlining the process for companies like Incarna and BlockRun. This service reduces integration time significantly, enabling agents to manage microtransactions efficiently with enforced spending limits.
Key evidence
- Incarna integrated AgentCore payments in three days, reducing the expected integration time from two to three months to just three days.
- BlockRun serves over 90 models from more than 15 providers, with each inference call quoted and settled independently.
- AgentCore payments enforces spending limits at the infrastructure level, preventing agents from exceeding their budget even if their prompts are manipulated.
Why it matters
The ability for AI agents to autonomously manage payments for services is crucial for scaling operations in AI applications. This capability allows for more efficient resource utilization and cost management, which is essential for developers looking to implement AI solutions in real-world scenarios. The integration of spending controls also enhances trust in autonomous systems, as it mitigates the risk of overspending.
📖 Reader Mode
~8 min readWhen an AI agent runs, it often needs to buy something to finish a task: a model inference, an API response, access web content, or a call to another agent. These purchases are small and frequent, sometimes a fraction of a cent each, and they happen inside the agent’s loop with no person available to approve them.
Amazon Bedrock AgentCore payments removes that burden. It gives agents a managed way to pay for services on demand, with spending limits enforced by the infrastructure rather than by the model. In this post, we look at how Incarna used AgentCore payments to let its agents pay BlockRun for model inference one request at a time. BlockRun is a pay-as-you-go inference router serving more than 90 models from more than 15 providers over x402, each call quoted and settled independently. AgentCore payments works with x402-compatible endpoints, including Amazon Bedrock inference endpoints. With the service, the Incarna team cut the work of adding x402 payment support from months to days, and put an end-to-end pay-per-inference flow into production.
The challenge: Paying for inference by the request
Paying per inference is a high-frequency, low-value pattern. An agent might make hundreds of small purchases in a single session, each worth a fraction of a cent. Card networks weren’t built for sub-cent payments. Building your own rails means solving several hard problems at once. You must decide where the money is held and how each payment is signed, support emerging payment protocols such as x402, and keep an autonomous agent from overspending.
What AgentCore payments provides
Amazon Bedrock AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model. AgentCore payments is a managed capability of Amazon Bedrock AgentCore that builders can use to add payments to their agents in a few lines of code. It handles the payment protocol, connects to a wallet, signs the transaction, and enforces spending limits. Builders integrate through one managed service instead of assembling the parts on their own. A few things make it a good fit for pay-per-inference:
- Managed wallets. Incarna provisions each agent’s wallet using the Coinbase CDP connector. The customer owns it and grants Incarna a delegated authorization to use the wallet.
- Native protocol handling. When a paid endpoint answers with HTTP 402 (“Payment Required”), the agent uses AgentCore payments to make the payment over x402. It signs the transaction with the configured wallet and returns cryptographic proof to the merchant.
- Spending governance at the infrastructure layer. AgentCore payments enforces limits outside the model, so an agent can’t exceed them even if its prompt is manipulated.
- Settlement you can audit. Payments settle in a stablecoin. Incarna uses USDC on the Base network, and each transaction is verifiable on-chain.
The following diagram illustrates the end-to-end architecture. AgentCore runs the agent, BlockRun serves metered inference, and AgentCore payments connects to the customer’s wallet, enforces the spending limit, and signs each payment on behalf of the agent’s Incarna identity.
Figure 1: Pay-per-inference architecture. Agent request flows through AgentCore to BlockRun (HTTP 402), with AgentCore payments signing the x402 payment from the agent’s own wallet
Source: incarna.io/aws-blockrun-incarna-partnership
Prerequisites
The path Incarna followed is open to other teams, and AgentCore payments provisions the pieces for you. You can set them up through a guided conversation with the AgentCore payments skill in the Agent Toolkit for AWS (in Claude Code, Kiro, or Codex). You can also create each one yourself with the AgentCore CLI, the AgentCore SDK, or the AWS SDK.
Start by storing your Coinbase CDP or Stripe Privy credentials as a payment credential provider, which keeps the secrets in AWS Secrets Manager instead of your code. Create a Payment Manager and connector to coordinate payments against those credentials, and set a default spending limit while you’re there. Then create a payment instrument, the embedded wallet the agent pays from. The end user funds it and grants signing permission through a redirect URL, and on a test network you can fund it with testnet USDC.
The flow: Buying and selling one inference
In this integration, BlockRun is the seller. BlockRun serves metered model inference from a live catalog, and each call is quoted, paid, and settled on its own. The flow for a single inference is straightforward:
- The agent needs a model call. It integrates with BlockRun, which handles provider selection and delivery. No per-provider subscription required.
- BlockRun answers with PaymentRequired challenge and a price for that specific call.
- It opens a payment session and calls
ProcessPayment. AgentCore payments checks the quote against the spending limits set for the session and signs the authorization from the agent’s own wallet address. - The seller verifies the payment signature.
- BlockRun serves the inference and records the charge, a small per-call amount.
Because settlement happens per request, the agent pays only for what it uses, and a call the agent chooses not to make costs nothing.
AgentCore payments supports two x402 payment schemes: exact and upto. The exact scheme is usually used when the price is known up front. The upto scheme suits resources with dynamic pricing. The agent authorizes a ceiling amount, and the inference provider settles for actual usage at the end, up to that ceiling.
Keeping spend under control
The piece that makes builders comfortable letting an agent move real money is the payment session. A session sets a ceiling, the most the agent can spend, and AgentCore payments enforces that ceiling at the infrastructure layer. The agent’s own code and prompt can’t change it. Each session also carries an expiry time, and Incarna sizes its sessions to a day’s budget. If something goes wrong in the agent’s logic, it still can’t spend past the limit the customer set. Even integrations that don’t use session-level budgets today can adopt it without code changes. For more on how these controls work, see Enable safe agentic payments with built-in guardrails.
Results
Incarna put the pay-per-inference flow into production on Base, with BlockRun serving the sell side and AgentCore payments governing every transaction.
“The next generation of AI agents shouldn’t have to choose between better performance and sustainable economics. BlockRun’s open source router gives developers full control over their own model set, while continuously benefiting from BlockRun’s ongoing benchmark-driven routing improvements—delivering higher task success rates at lower token cost. And with AgentCore payments and x402 providing the spending controls agents need, that intelligence can be put to work safely in the real world.”
— Vicky Fu, Founder, BlockRun
The Incarna team completed the full AgentCore payments integration in three days: one to build and two to test. That was roughly 200 lines of application code, against the two to three months originally scoped. Across the beta, agents have processed over 1,000 payments ranging from $0.001 to $0.05 per call, each settling individually on-chain.
“AgentCore payments covered everything we needed for an agent to pay over x402: a wallet the customer owns, a funding and revocation flow, a spending limit the platform enforces, and signing that handles both versions of x402. We wrote none of it.”
— Justin Zhou, Founder, Incarna
Conclusion
AgentCore payments gives agents a governed, on-demand way to pay for the services they use. BlockRun and Incarna show how it comes together end to end: an agent that pays for inference, a provider that meters and settles it, and an identity that owns the transaction. If you’re building agents that need to buy services as they run, you can add that capability through a single, governed integration point with AgentCore payments.
Getting started
Ready to add pay-per-inference to your agents? Here’s how to start:
- Set up AgentCore payments. Create a Payment Manager with your wallet connection and spending policies. Connect a Coinbase CDP wallet as your credential provider. See the AgentCore payments developer guide for step-by-step instructions.
- Open a payment session with your budget. Before the agent starts a task, open a session with a spending cap that fits the workload. The agent transacts within that ceiling and the infrastructure enforces it.
- Call a paid endpoint. Point your agent at an x402-compatible service (like BlockRun). When the endpoint returns HTTP 402, call
ProcessPaymentwith the payment details. AgentCore payments handles the signing and returns proof the agent can present to access the service.
To explore the code, see the AgentCore payments samples on GitHub.
Learn more
- Amazon Bedrock AgentCore payments is now generally available: Enabling agents to transact safely and autonomously at scale
- Technical deep dive: AgentCore payments and innovation in agentic commerce
- Enable safe agentic payments with built-in guardrails using AgentCore payments
About BlockRun and Incarna
BlockRun is a pay-as-you-go inference router that serves model inference over the x402 payment protocol. Agents reach a live catalog of models through a single metered endpoint. Each call is independently quoted, authorized, paid, and settled on Base (USDC). BlockRun handles provider selection and delivery so agents access many models through one integration without per-provider subscriptions or billing relationships.
Incarna, built by SpreadX, gives AI agents a persistent identity that survives sessions, models, and runtimes. Each identity carries its own wallet, email address, social accounts, and an action history that stays attached to one ID across runs. When an agent pays for a service, the on-chain payer is the agent’s own identity rather than a shared platform key, making every transaction attributable to a single durable entity. Incarna is live on Base mainnet with real settlement.
About the authors
— Originally published at aws.amazon.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from AWS Machine Learning
See more →
Building an agentic app deployer with Amazon Bedrock and AWS Lambda
PDI Technologies developed PDI Brew, enabling non-technical employees to create web applications on AWS without developer involvement, leveraging Amazon Bedrock for AI capabilities. This agentic app deployer streamlines internal tool delivery, removing traditional bottlenecks in deployment pipelines.


