My demo agent pays a vendor, then tries to withdraw from a vault. The trace shows two green tool calls.
One of those transactions reverted.
TL;DR
- Agent traces end at the tool call. What happened on chain (mined, reverted, replaced, what it cost) is not in them.
- Key the transaction by its hash and record it as two spans inside the tool span:
sendandconfirm.- The hard parts are replaced transactions, the L1 data fee and revert reasons. Here’s how each one is handled.
What does the trace actually show?
Take an AI SDK agent with two tools: pay_vendor and withdraw_from_vault. Each tool sends a transaction with viem.
Without extra instrumentation, the trace looks like this:
invoke_agent treasury-agent├─ execute_tool pay_vendor ✓└─ execute_tool withdraw_from_vault ✓Both tools returned, so both spans are fine. The tool did its job: it sent a transaction and got a hash back.
But the hash is where the trace stops. Was the transaction mined? Did it revert? How much did it cost? That lives in a block explorer, in another tab, joined by copy-pasting a hash.
Why does this matter more for agents?
A human who sends a transaction watches the wallet. An agent doesn’t. It moves to the next step.
More agents now have wallets. Standards like x402 describe payments between clients and servers that agents can make on their own. When an agent spends, “the tool returned” and “the payment went through” are two different facts.
Debugging an agent means reading its decisions next to their results. AI agent observability tools built on OpenTelemetry show the decisions. The results are in another system.
Where does the chain outcome live today?
It usually ends up in one of these places. None of them is the agent’s trace by default:
| Where | What you get | What’s missing |
|---|---|---|
| Block explorer | Status, gas, logs for one hash | Which agent, which step, which prompt |
| Contract / address monitoring | Alerts on a contract or wallet | The agent run that caused it |
| Agent observability tool | The LLM call and the tool call | The transaction outcome |
Each one answers part of the question. The join key between them is the transaction hash.
What if the hash were the key?
That’s the idea behind hashspan. Every transaction becomes two OpenTelemetry spans, nested under the tool that sent it:
send {chainId}ends when the hash comes back, or when sending fails.confirm {chainId}ends when the receipt arrives, or on timeout. It links back to itssendspan.
Why two spans and not one? Sending and confirming are different events: a send can succeed while its confirmation reverts, times out or gets replaced, and the receipt is often awaited somewhere else in the code, or not at all.
The same agent run now looks like this. Everything in this post comes from the example agent, which runs on a local Anvil chain (chain id 31337):
invoke_agent gen_ai.agent.name=treasury-agent├─ step 1│ └─ execute_tool pay_vendor│ ├─ send 31337 blockchain.tx.value=250000000000000000│ └─ confirm 31337 blockchain.tx.status=success└─ step 2 └─ execute_tool withdraw_from_vault ├─ send 31337 blockchain.contract.function.name=withdraw └─ confirm 31337 blockchain.tx.status=reverted blockchain.tx.revert.reason=WithdrawalLimitExceeded(100000000000000000, 1000000000000000000)The tools contain no tracing code. The AI SDK runs each tool inside its execute_tool span, so the transaction spans become its children.

make demo run. Click for the full trace.These are standard spans. Jaeger, Grafana Tempo, Langfuse or Honeycomb show them like any other.
What goes on the confirm span?
The confirm span carries the outcome:
- blockchain.tx.status
- reverted
- blockchain.block.number
- 2
- blockchain.tx.gas.used
- 21,660
- blockchain.tx.fee
- 40,616,290,500,000
- blockchain.tx.revert.reason
- WithdrawalLimitExceeded(100000000000000000, 1000000000000000000)
- gen_ai.agent.name
- treasury-agent
| Attribute | Meaning |
|---|---|
blockchain.tx.status |
success, reverted or replaced. A confirm that gave up waiting has error status and error.type timeout |
blockchain.block.number |
The block it landed in |
blockchain.tx.gas.used |
Gas used |
blockchain.tx.fee |
Gas used × effective gas price + L1 data fee, in wei |
blockchain.tx.revert.reason |
The decoded reason, when the transaction reverted |
Agent identity rides along as gen_ai.agent.id and gen_ai.agent.name, from the OpenTelemetry GenAI semantic conventions (in their own repository since June 2026, still in Development). So your backend can answer “which agent spent what” without joining anything.
Which parts are harder than they look?
Recording a hash and a receipt sounds like twenty lines of code. I thought so too. These three cases took most of the work.
1. The receipt isn’t for your hash
A pending transaction can be replaced by another one with the same sender and nonce: sped up, or cancelled. viem’s waitForTransactionReceipt then resolves with the receipt of the replacing transaction.
If you record that receipt on the span of the hash you waited for, the fee and status belong to the wrong transaction. hashspan ends the waited-for span as replaced, with blockchain.tx.replacement.hash, and puts the receipt on the span of the hash that was mined.
2. The fee is missing a part
On OP-stack chains such as Base, the total cost has two components (the local demo chain has no L1 fee, so this one isn’t in the screenshot): the execution gas fee and the L1 data fee (Optimism docs). Gas used × gas price leaves the second one out.
blockchain.tx.fee includes it, and blockchain.tx.l1_fee records it on its own.
3. “Reverted” doesn’t say why
A receipt only tells you the transaction reverted. The reason isn’t stored on chain.
hashspan replays the reverted transaction with eth_call against the state of the previous block, then decodes the revert data: an Error(string) message, a Panic code, or a custom error like WithdrawalLimitExceeded(limit, requested) when the ABI is known.
This is best effort. The replay doesn’t see transactions that ran earlier in the same block. If one of them changed the state yours depended on, the reason can be missing or different. Some RPC providers also don’t serve historical state.
One catch: the confirm span of a revert ends after the replay, so it can outlive the tool call. Short-lived scripts and serverless functions should call flush() before shutting down, or the most interesting span never gets exported.
What doesn’t this do?
It’s a tracing library, so a few things are out of scope:
- No alerting. It records what happened; it doesn’t page you. Pair it with a monitoring tool for that.
- No signing or sending. It observes your transactions. It never signs or broadcasts.
- No LLM evals. Prompt and response tracing stays with your agent observability tool. hashspan adds spans to the same trace.
Recording sensitive data is opt-in. Calldata arguments and error messages are off until you enable them, and addresses can be hashed or dropped. The hash still resolves to the parties on chain, so this limits what your backend stores; it isn’t anonymity.
Try it in two commands
The example runs on a local chain and needs no API key. In a clone of the repository, with Docker running and Node.js 22.18 or later (see the local lab):
make lab-up # Jaeger on http://localhost:16686make demo # runs the agent, including the reverted withdrawalOpen Jaeger, pick treasury-agent, and expand withdraw_from_vault.
The span and attribute names are still a draft. If you trace agents that send transactions, tell me what’s missing.
Discussion
Comments live in GitHub Discussions. Sign in with GitHub to join.